DSDyn-VLA: A Dual-Stream Dynamic Manipulation Framework with Motion Perception, Future Awareness, and Realtime Correction

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bottlenecks of perception latency and control lag in Vision-Language-Action (VLA) models operating within dynamic environments by proposing a dual-stream dynamic framework. The proposed method integrates optical flow estimation for motion planning with reinforcement learning policies to enable real-time residual correction, thereby establishing a mechanism for highly consistent action prediction and closed-loop rectification. Experimental evaluations on DynBench, a MuJoCo-based simulation benchmark, demonstrate that this framework reduces the Kinetix failure rate by 76%. Furthermore, it achieves a sixfold improvement in real-world manipulation success rate over PI0.5 and yields a fivefold enhancement in simulated performance, significantly bolstering the robustness of VLA models when manipulating moving objects.
📝 Abstract
While Vision-Language-Action (VLA) models excel in static tasks, they struggle in dynamic environments where objects are in motion (e.g., conveyor belt manipulation). We identify three fundamental limitations hindering current VLAs in these scenarios: the \textbf{perception gap}, where static visual inputs lack temporal motion cues; the \textbf{latency gap}, where inference delays render actions obsolete; and the \textbf{control gap}, caused by the open-loop action chunk execution without real-time adjustment. In this work, we propose \textbf{DSDyn-VLA}, a Slow-Fast \textbf{D}ual-\textbf{S}tream \textbf{Dyn}amic manipulation framework that integrates motion-aware foresighted planning with real-time residual correction. The slow \textbf{Flow-Planner} serves as a macro-planner. By enhancing the VLA with optical flow for temporal perception and a future state awareness mechanism to preemptively offset inference latency, it produces globally consistent, motion-aware action chunks. Complementing this, the fast \textbf{Res-Refiner} employs a lightweight RL policy to inject high-frequency, closed-loop corrections into the planned action chunks based on real-time observations. In addition, we introduce \textbf{DynBench}, a MuJoCo-based benchmark for dynamic object manipulation that comprises nine tasks. Extensive experiments demonstrate that DSDyn-VLA reduces the failure rate by over 76\% compared to current SOTA method in high-latency setting on the Kinetix dynamic benchmark, while achieving about 6$\times$ the success rate of PI0.5 in real-world dynamic settings and about 5$\times$ on DynBench. We will open-source all the code and weights.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
dynamic manipulation
perception gap
inference latency
real-time correction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action
Dual-Stream Framework
Optical Flow
Real-time Correction
Dynamic Manipulation
💼 Related Jobs
No related jobs found.
W
Wenhao Li
University of Sydney
X
Xiu Su
Central South University
Y
Yu Han
University of California, San Diego
Y
Yichao Cao
Central South University
Shan You
Shan You
SenseTime Research
deep learningmultimodal LLMedge AI
Chang Xu
Chang Xu
University of Sydney
Machine LearningComputer Vision and MultimediaData Mining