REACT: Rolling Denoising and Dual Decoupling for Reactive Robot Control with VLA Models

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent trade-off between trajectory smoothness and real-time responsiveness caused by action chunking in streaming Vision-Language-Action (VLA) models. To this end, we propose a rolling denoising framework coupled with a dual-decoupled architecture. The former maintains a persistent action buffer with staggered timesteps, progressively refining future actions using the latest observations rather than generating chunks from scratch. The latter disentangles perceptual encoding from execution modules to enable parallel computation, thereby enhancing real-time processing capabilities. Built upon a diffusion Transformer, our streaming VLA model significantly reduces reaction latency, improves task success rates, and yields smoother trajectories in both the RoboTwin 2.0 simulation and real-world bimanual manipulation tasks, effectively resolving the coherence-reactivity dilemma in robot control.
📝 Abstract
Flow-based vision-language-action (VLA) models generate action chunks for temporally coherent robot motion, but chunked control creates a fundamental closed-loop trade-off: long chunks provide smooth execution, whereas frequent replanning improves reactivity at the cost of action discontinuities. We introduce REACT, a rolling-denoising framework that makes flow-based VLAs more reactive while preserving long-horizon context. Instead of regenerating entire action chunks from scratch, REACT maintains a persistent action buffer with staggered flow timesteps. At each control step, the full horizon is denoised using the latest observation, the cleanest action block is executed, partially refined future blocks are shifted forward, and fresh noise is appended to the tail. As a result, each executed action block is refined across multiple recent observations before deployment. To support real-time control, we further introduce dual decoupling, which separates sensing, VLM encoding, DiT denoising, and action execution, enabling high-frequency observation updates and action streaming under practical compute constraints. Across the RoboTwin 2.0 simulation benchmark and real-world tasks spanning bimanual manipulation and dynamic control on multiple robot platforms, REACT improves task success and reduces reaction latency while producing smoother trajectories than frequent-replanning and asynchronous baselines.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action Models
Reactive Robot Control
Action Chunking
Closed-loop Trade-off
Real-time Control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Rolling Denoising
Dual Decoupling
Vision-Language-Action Models
Flow Matching
Reactive Robot Control
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Houlong Xiong
Shanghai Jiao Tong University
Z
Zhenqi Qiu
PrimeBot
Z
Zechen Wang
PrimeBot
S
Suohang Zhang
Zhejiang University
Y
Yiyu Ren
PrimeBot
Wanting Xu
Wanting Xu
ShanghaiTech University
Computer VisionRobotics
H
Hongfei Niu
PrimeBot
C
Chengyang He
National University of Singapore
Ge Sun
Ge Sun
Research Hydrologist and Project Leader, USDA Forest Service
Forest HydrologyEcohydrologyWatershed Management
R
Ran Cheng
PrimeBot
Q
Qian Zhu
PrimeBot