stream processing

Designs, builds, and analyzes systems and pipelines that continuously ingest, transform, and route unbounded sequences of records or events for low‑latency and/or real‑time processing. Work includes implementing stream operators (filters, windowed aggregations, joins), managing state and time semantics, and ensuring correctness, fault tolerance, and scalable delivery guarantees for continuous data flows.

streamprocessing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.84
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$196K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Stream data processing faces persistent challenges including manual incremental maintenance, SQL semantic inconsistencies between streaming and batch paradigms, and insufficient enterprise-grade operational capabilities; moreover, existing systems overemphasize sub-second latency, limiting applicability to mainstream second- to minute-scale analytical workloads. To address these, we propose Delayed View Semantics (DVS), the first formal framework unifying stream and batch semantics across a broad latency spectrum. We design declarative dynamic table primitives to ensure end-to-end transactional consistency and high availability, and extend the snapshot isolation model to enforce invariants in streaming applications. Evaluated on a production system built atop Snowflake, our approach reduces development complexity by over 80%, supports configurable end-to-end latency from milliseconds to minutes, cuts operational overhead by 90%, and currently serves thousands of enterprise customers.

Addresses challenges in building streaming data pipelinesOptimizes for diverse latency requirements in streamingResolves semantic gaps between streaming and databases

Existing LLM frameworks are stateless and execute queries in isolation, making them ill-suited for long-horizon, semantics-aware analysis of unstructured data streams. Method: This paper introduces the first LLM-driven continuous stream processing paradigm, extending Retrieval-Augmented Generation (RAG) to streaming settings and formalizing *continuous semantic operators*. We design a dynamic optimization framework integrating lightweight shadow execution with multi-objective Bayesian optimization (MOBO) to adaptively balance throughput and accuracy. Further, we incorporate tuple batching, operator fusion, and cost-aware scheduling within the VectraFlow system. Results: Experiments demonstrate that VectraFlow responds to load fluctuations in real time, sustains high-accuracy continuous semantic querying under high throughput, achieves significant efficiency gains, and incurs only bounded, controllable accuracy degradation.

Enables LLM reasoning for continuous unstructured stream processingIntroduces optimizations balancing accuracy and efficiency in stream analyticsSupports persistent semantic queries over evolving unstructured data streams

To address insufficient multi-objective load balancing in stream processing systems under complex workloads, this paper proposes a multi-tier collaborative scheduling framework. The framework introduces dynamic inter-layer coordination mechanisms and lightweight interfaces among schedulers, enabling seamless integration of novel scheduling policies. It jointly optimizes computational resource utilization, end-to-end latency, and throughput by integrating multi-objective optimization, distributed resource management, and real-time feedback control. Its key innovation lies in shifting hierarchical scheduling from static decoupling to dynamic collaboration—preserving scalability while significantly enhancing adaptability. Evaluated in Meta’s production environment, the system reliably processes TB-scale data with sub-second latency; it improves critical resource utilization by 27% and reduces tail latency by 41%.

Designing co-operation in hierarchical multi-objective schedulers for stream processingEnhancing load balancing across compute resources for growing application complexityIntegrating new schedulers into existing hierarchies to improve proactive resource management

Towards Fine-Grained Scalability for Stateful Stream Processing Systems

Mar 14, 2025
YQ
Yunfan Qing
🏛️ Shanghai Jiao Tong University

Existing stateful stream processing systems suffer from high latency, processing pauses, and even service outages during dynamic scaling due to coarse-grained synchronization and inefficient state migration. This paper proposes DRRS, a novel scaling approach that introduces fine-grained scaling signals and data rerouting to enable record-level deterministic scheduling—thereby eliminating processing suspension—and employs sub-scale state partitioning with sharded state migration to minimize dependency overhead. Implemented atop Apache Flink, DRRS supports real-time trigger and seamless transition. Experimental evaluation demonstrates that, compared to state-of-the-art methods, DRRS reduces peak and average latency by 81.1% and 95.5%, respectively, shortens scaling time by 72.8%–86%, and incurs zero interruption during non-scaling periods. These results significantly enhance the real-time responsiveness and reliability of elastic scaling in stateful stream processing.

Addresses performance degradation in stream processing systems.Improves state migration and synchronization for better runtime adaptability.Reduces latency and scaling duration in dynamic scaling.

Cloud application development has long suffered from high expertise barriers due to the need to integrate distributed systems, database, and software engineering knowledge. This paper proposes Stateflow—the first cloud-native, streaming, stateful function-computing framework—designed to jointly address programmability, strongly consistent fault-tolerant transactions, and serverless semantics. Its three key contributions are: (1) an object-oriented declarative programming model that eliminates explicit error handling; (2) the Styx engine, which guarantees deterministic multi-partition serializability, snapshot consistency, and zero-loss state migration; and (3) a streaming dataflow execution model with transaction-aware state migration, enabling dynamic elastic scaling. Experimental evaluation demonstrates that Stateflow significantly outperforms existing systems in both throughput and recovery performance, substantially reducing development complexity for high-concurrency, strongly consistent cloud applications.

Democratizing scalable cloud application developmentEnabling serverless semantics for streaming dataflowsProviding high-performance fault-tolerant serializable transactions

Latest Papers

What's happening recently
View more

This work addresses the high latency in real-time stream processing caused by tight coupling between state I/O and the data path, which blocks the CPU on the critical path. To mitigate this, the authors propose Keyed Prefetching, a mechanism that extracts state-access keys from upstream operators and proactively prefetches the required state, thereby overlapping I/O with computation to hide latency. Complementing this, a Timestamp-Aware Caching strategy is introduced to efficiently manage prefetched and historical states in memory. Together, these techniques significantly reduce end-to-end latency for long-running real-time queries while maintaining high throughput and effectively decoupling state access from data processing.

I/O latencylow-latencyreal-time queries

Existing evaluation approaches for streaming process mining algorithms predominantly rely on static logs or synthetic event streams, which fail to capture the complexity of real-world event streams in IoT environments—such as out-of-order events, concurrency, incomplete cases, and concept drift. This work addresses this gap by introducing, for the first time, a feature framework from data stream research into streaming process mining. It proposes an intent-oriented event stream generation methodology, extends the conceptual model of event streams, and implements a prototype tool, Stream of Intent. This tool enables customizable configuration of key stream characteristics reflective of real-world scenarios, facilitating the generation of controlled, reproducible, and realistically complex event streams. Consequently, it significantly enhances the relevance and adaptability of algorithm evaluation and development in streaming process mining.

BenchmarkingConcept DriftEvent Streams

Existing LLM streaming frameworks lack unified support for token-level flow control, queue management, scheduling, and backpressure, often relying on ad hoc callbacks that lead to system complexity and unreliability. This work proposes AiFlow—the first token-native reactive orchestration model—which formalizes LLM generation streams as typed Context<T> events propagated through a directed stream graph. Node guardians uniformly manage queue boundaries, concurrency, ordering, and fault tolerance. AiFlow supports DSL/JSON graph compilation and provides static verification for type safety, stateful concurrency, and injection compatibility. Its reactive architecture, combined with runtime-enforced policies, guarantees bounded memory usage. Experiments demonstrate a 70.9–94.7% reduction in time-to-first-token latency, stable and controllable queue depths, and a 93.7–96.5% decrease in maximum queue length.

backpressurequeue managementreactive orchestration

Hot Scholars

BF

Bernd Finkbeiner

Professor of Computer Science, CISPA Helmholtz Center for Information Security
Reactive SystemsVerificationSynthesisTemporal Logic
ML

Mian Lu

4Paradigm Technology
machine learning systemsGPGPUhigh performance computing
XZ

Xuanhe Zhou

Assistant Professor, Shanghai Jiao Tong University
Data ManagementArtificial Intelligence
GL

Guoliang Li

Professor, Tsinghua University
DatabaseBig DataCrowdsourcingData Cleaning & Integration