🤖 AI Summary
This study addresses the action distribution shift and limited real-time reactivity in robotic policies caused by asynchronous execution. To mitigate these issues, we propose a dual mechanism comprising recursive flow field distillation and a proposal-parsing framework. By leveraging asynchronous multi-sequence generation and likelihood-based action selection, our method aligns the asynchronous action distribution with that of the original Vision-Language-Action (VLA) model, effectively rectifying distributional biases in non-Markovian demonstrations and restoring control reactivity. Experimental evaluations demonstrate that the proposed approach matches the success rate of the original VLA on the LIBERO benchmark while maintaining approximately 80% success on RoboMimic—representing a 30-percentage-point improvement over existing methods—and achieves near-complete coverage of the action space.
📝 Abstract
Generalist robot policies such as vision-language-action models (VLAs) have achieved remarkable generalization, but their inference delays can conflict with the demands of real-time control. Asynchronous execution avoids pauses between action chunks by predicting the next sequence of actions while the robot carries out the previous one. In this paper, we study whether asynchronous execution produces the same action distribution as the original VLA. We find that, for non-Markovian demonstrations, asynchronous execution can produce a fundamentally different action distribution, which can limit the policy's reactivity. In our method, we seek to restore this reactivity by aligning the asynchronously produced action distribution with that of the original VLA through two complementary mechanisms. First, Recursive Flow-Field Distillation trains the asynchronous policy using the VLA's action-generation flow. We characterize the learned distribution theoretically and show experimentally that our asynchronous policy can generate nearly the full range of actions the original VLA would produce, while existing asynchronous methods recover only a fraction of that range. Second, Propose-Resolve prepares multiple action sequences asynchronously and uses the latest observation to select among them based on a lightweight approximation of their likelihood under the VLA's action distribution. Our resulting method matches the original VLA's success on LIBERO and retains about 80% of its success on RoboMimic, about 30 percentage points more than existing asynchronous methods.