🤖 AI Summary
This study addresses the inherent trade-off between trajectory smoothness and real-time responsiveness caused by action chunking in streaming Vision-Language-Action (VLA) models. To this end, we propose a rolling denoising framework coupled with a dual-decoupled architecture. The former maintains a persistent action buffer with staggered timesteps, progressively refining future actions using the latest observations rather than generating chunks from scratch. The latter disentangles perceptual encoding from execution modules to enable parallel computation, thereby enhancing real-time processing capabilities. Built upon a diffusion Transformer, our streaming VLA model significantly reduces reaction latency, improves task success rates, and yields smoother trajectories in both the RoboTwin 2.0 simulation and real-world bimanual manipulation tasks, effectively resolving the coherence-reactivity dilemma in robot control.
📝 Abstract
Flow-based vision-language-action (VLA) models generate action chunks for temporally coherent robot motion, but chunked control creates a fundamental closed-loop trade-off: long chunks provide smooth execution, whereas frequent replanning improves reactivity at the cost of action discontinuities. We introduce REACT, a rolling-denoising framework that makes flow-based VLAs more reactive while preserving long-horizon context. Instead of regenerating entire action chunks from scratch, REACT maintains a persistent action buffer with staggered flow timesteps. At each control step, the full horizon is denoised using the latest observation, the cleanest action block is executed, partially refined future blocks are shifted forward, and fresh noise is appended to the tail. As a result, each executed action block is refined across multiple recent observations before deployment. To support real-time control, we further introduce dual decoupling, which separates sensing, VLM encoding, DiT denoising, and action execution, enabling high-frequency observation updates and action streaming under practical compute constraints. Across the RoboTwin 2.0 simulation benchmark and real-world tasks spanning bimanual manipulation and dynamic control on multiple robot platforms, REACT improves task success and reduces reaction latency while producing smoother trajectories than frequent-replanning and asynchronous baselines.