Score
Designs and implements end-to-end systems and components that ingest, route, process, and store continuous data streams with attention to low latency, scalability, fault tolerance, and memory efficiency. Work includes architecting streaming computations and pipelines, designing streaming data structures and chunk-wise/causal processing protocols, implementing lookahead and dual-buffer strategies, selective and online state updates, out-of-core processing, and orchestration of real-time streaming inference and adaptation.
To address insufficient multi-objective load balancing in stream processing systems under complex workloads, this paper proposes a multi-tier collaborative scheduling framework. The framework introduces dynamic inter-layer coordination mechanisms and lightweight interfaces among schedulers, enabling seamless integration of novel scheduling policies. It jointly optimizes computational resource utilization, end-to-end latency, and throughput by integrating multi-objective optimization, distributed resource management, and real-time feedback control. Its key innovation lies in shifting hierarchical scheduling from static decoupling to dynamic collaboration—preserving scalability while significantly enhancing adaptability. Evaluated in Meta’s production environment, the system reliably processes TB-scale data with sub-second latency; it improves critical resource utilization by 27% and reduces tail latency by 41%.
This study addresses the low-latency, high-throughput, and scalable data transfer requirements for cross-facility (edge-to-HPC) workflows in AI–HPC convergence scenarios. Method: We systematically compare three streaming architectures—Direct Transfer Streaming (DTS), Proxy-based Streaming (PRS), and Management Service Streaming (MSS)—and propose the DS2HPC taxonomy alongside SciStream, a lightweight in-memory streaming toolkit supporting scientific workflow patterns including work sharing, feedback loops, and broadcast-aggregation. Contribution/Results: Evaluation on real-world advanced computing infrastructures shows that DTS achieves minimal latency and maximal throughput but suffers from deployment constraints; MSS offers strong scalability yet incurs significant overhead; PRS delivers near-DTS performance while maintaining deployment flexibility and scalability, making it the optimal trade-off between efficiency and practicality. Our work provides empirical evidence and methodological support for designing cross-domain scientific data streaming architectures.
This work addresses the inefficiency of existing serverless computing and stream processing systems in handling short-lived, lightweight, and unpredictable stateful data streams. To bridge this gap, the paper proposes “stream functions”—an extension to the function-as-a-service model that elevates short streams to first-class units of execution, state management, and autoscaling. Stream functions express inter-event logic through iterator-based interfaces, effectively integrating stream processing semantics with the elasticity of serverless architectures. This approach is the first to treat short streams as fundamental execution units, thereby filling a critical void in lightweight stateful stream processing. Experimental evaluation in video processing scenarios demonstrates that the proposed system reduces runtime overhead by approximately 99% compared to mainstream stream processing engines, while maintaining performance comparable to conventional serverless functions.
This work addresses the challenge of maintaining both timeliness and stability in real-time data streams within scientific workflows, which are highly susceptible to hardware failures, network disruptions, and performance fluctuations in complex environments. The authors propose a lightweight, non-intrusive fault-tolerance mechanism that integrates asynchronous, non-blocking checkpointing with a progress-aware dynamic load redistribution strategy. This approach enables efficient fault recovery and resource rebalancing without interrupting ongoing computations. Under fault-free conditions, the method incurs less than 1% runtime overhead, while in high-failure-rate scenarios, it reduces the impact of faults and performance anomalies by up to sixfold, substantially enhancing the resilience and resource utilization of stream processing systems.
This paper addresses the end-to-end Quality of Experience (QoE) assurance challenge for video streaming over best-effort networks. It systematically analyzes bottlenecks across the full pipeline—from video acquisition and compression (H.264/HEVC/AV1), upload, transcoding, CDN scheduling, adaptive bitrate (ABR) decision-making, to playback. We propose the first unified end-to-end pipeline analytical framework, classifying and modeling over 200 works along two orthogonal dimensions: methodology (heuristic, optimization, machine learning) and technical characteristics (codecs, super-resolution, etc.). The resulting methodology map is the most comprehensive to date, rigorously delineating performance boundaries and industrial deployment constraints for each approach. Our analysis identifies critical evolutionary trends—including ultra-low-latency live streaming, AI-native video coding, and edge-coordinated delivery—providing a systematic reference for both academic research and industry implementation.
This work addresses the inefficiency of manual performance tuning in cloud-native stream processing systems, which heavily relies on expert experience. To automate and accelerate configuration optimization, the authors propose an experiment-driven approach that integrates Latin hypercube sampling, simulated annealing, and hill climbing into a three-stage search strategy. This method is deeply coupled with the Theodolite benchmarking framework to automatically orchestrate experiments on Kubernetes and preemptively terminate underperforming configurations. Evaluated on Kafka Streams, the approach efficiently explores the configuration space and identifies settings that substantially outperform default configurations, achieving up to a 23% improvement in throughput. The study demonstrates a practical and effective pathway toward automated, high-efficiency tuning of stream processing systems in cloud-native environments.
This work addresses the high latency in real-time stream processing caused by tight coupling between state I/O and the data path, which blocks the CPU on the critical path. To mitigate this, the authors propose Keyed Prefetching, a mechanism that extracts state-access keys from upstream operators and proactively prefetches the required state, thereby overlapping I/O with computation to hide latency. Complementing this, a Timestamp-Aware Caching strategy is introduced to efficiently manage prefetched and historical states in memory. Together, these techniques significantly reduce end-to-end latency for long-running real-time queries while maintaining high throughput and effectively decoupling state access from data processing.
This work addresses the challenges of resource contention and performance instability in edge stream processing under dynamic workloads and fluctuating resources, which often stem from the absence of a unified coordination mechanism. To this end, we propose the first framework that integrates embodied intelligent agents into collaborative optimization for edge stream processing. Our approach features a context-aware autoscaling platform that unifies service-specific policies with global resource scheduling through an extensible monitoring interface and a multi-service action space exploration mechanism. Leveraging reinforcement learning–driven agents, the framework significantly enhances resource efficiency and response timeliness on real-world edge platforms, while also supporting user-defined policies and enabling reproducible experiments with visual analytics.
This work addresses the challenges of slow failure recovery, poor stability, and high operational costs that Apache Flink faces in large-scale production environments, which hinder its ability to meet stringent service-level objectives (SLOs). To overcome these limitations, the authors propose the first systematic approach that integrates engine-level and cluster-level elasticity. The solution innovatively combines runtime optimizations, fine-grained fault tolerance, a hybrid replication strategy, and high-availability mechanisms leveraging external systems, complemented by a highly reliable automated testing and deployment pipeline. Evaluated on ByteDance’s ultra-large-scale Flink clusters, the proposed framework significantly enhances system elasticity, stability, and recovery efficiency, thereby effectively ensuring compliance with demanding SLOs.