🤖 AI Summary
Existing stateful stream processing systems suffer from high latency, processing pauses, and even service outages during dynamic scaling due to coarse-grained synchronization and inefficient state migration. This paper proposes DRRS, a novel scaling approach that introduces fine-grained scaling signals and data rerouting to enable record-level deterministic scheduling—thereby eliminating processing suspension—and employs sub-scale state partitioning with sharded state migration to minimize dependency overhead. Implemented atop Apache Flink, DRRS supports real-time trigger and seamless transition. Experimental evaluation demonstrates that, compared to state-of-the-art methods, DRRS reduces peak and average latency by 81.1% and 95.5%, respectively, shortens scaling time by 72.8%–86%, and incurs zero interruption during non-scaling periods. These results significantly enhance the real-time responsiveness and reliability of elastic scaling in stateful stream processing.
📝 Abstract
Dynamic scaling is critical to stream processing engines, as their long-running nature demands adaptive resource management. Existing scaling approaches easily cause performance degradation due to coarse-grained synchronization and inefficient state migration, resulting in system halt or high processing latency. In this paper, we propose DRRS, an on-the-fly scaling method that reduces performance overhead at the system level with three key innovations: (i) fine-grained scaling signals coupled with a re-routing mechanism that significantly mitigates propagation delay, (ii) a sophisticated record-scheduling mechanism that substantially reduces processing suspension, and (iii) subscale division, a mechanism that partitions migrating states into independent subsets, thereby reducing dependency-related overhead to enable finer-grained control and better runtime adaptability during scaling. DRRS is implemented on Apache Flink and, when compared to state-of-the-art approaches, reduces peak and average latencies by up to 81.1% and 95.5% respectively, while achieving a 72.8%-86% reduction in scaling duration, without disruption in non-scaling periods.