🤖 AI Summary
This study addresses the limitation of existing Vision-Language-Action (VLA) models whose action representations neglect visual dynamics, proposing a visually grounded method for constructing an action latent space. The core innovation lies in explicitly binding continuous action latents to future scene dynamics. By training an action variational autoencoder (VAE), the proposed approach aligns the action space with visual prediction, while introducing a plug-and-play interface compatible with various mainstream architectures. Experimental results demonstrate that this method achieves a 98.1% success rate on the LIBERO benchmark and yields substantial performance improvements in both RoboTwin simulations and real-world bimanual robotic tasks.
📝 Abstract
Vision-Language-Action (VLA) models have become a central paradigm for robot policy learning, which predict actions in three forms: raw action chunks, discrete action tokens, or continuous action latents. However, existing action representations primarily model action trajectories, with limited consideration of the visual dynamics induced by these actions. We introduce ViDAL, a Visual Dynamics-grounded Action Latent Space that anchors continuous action latents in the future visual dynamics of the scene. Specifically, ViDAL learns action latent space by training an Action Variational Autoencoder (Action VAE) to reconstruct action chunks while aligning its latent with future scene dynamics. When integrated into downstream robot policies, the proposed Action VAE serves as a plug-in action interface compatible with multiple VLA architectures and enables optional future-video prediction as an additional capability. Empirically, ViDAL outperforms competitive baselines on LIBERO with 98.1% average success, improves a multi-task $π_{0.5}$ policy on RoboTwin 2.0 from 54.3% to 65.5% (Clean) and from 33.2% to 43.1% (Random) success rates over 50 dual-arm tasks, and yields 20.0% and 23.4% absolute success-rate gains on real-world single-arm Franka and dual-arm ARX robot platforms.