ViDAL: A Visual Dynamics-Grounded Action Latent Space for Vision-Language-Action Models

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing Vision-Language-Action (VLA) models whose action representations neglect visual dynamics, proposing a visually grounded method for constructing an action latent space. The core innovation lies in explicitly binding continuous action latents to future scene dynamics. By training an action variational autoencoder (VAE), the proposed approach aligns the action space with visual prediction, while introducing a plug-and-play interface compatible with various mainstream architectures. Experimental results demonstrate that this method achieves a 98.1% success rate on the LIBERO benchmark and yields substantial performance improvements in both RoboTwin simulations and real-world bimanual robotic tasks.
📝 Abstract
Vision-Language-Action (VLA) models have become a central paradigm for robot policy learning, which predict actions in three forms: raw action chunks, discrete action tokens, or continuous action latents. However, existing action representations primarily model action trajectories, with limited consideration of the visual dynamics induced by these actions. We introduce ViDAL, a Visual Dynamics-grounded Action Latent Space that anchors continuous action latents in the future visual dynamics of the scene. Specifically, ViDAL learns action latent space by training an Action Variational Autoencoder (Action VAE) to reconstruct action chunks while aligning its latent with future scene dynamics. When integrated into downstream robot policies, the proposed Action VAE serves as a plug-in action interface compatible with multiple VLA architectures and enables optional future-video prediction as an additional capability. Empirically, ViDAL outperforms competitive baselines on LIBERO with 98.1% average success, improves a multi-task $π_{0.5}$ policy on RoboTwin 2.0 from 54.3% to 65.5% (Clean) and from 33.2% to 43.1% (Random) success rates over 50 dual-arm tasks, and yields 20.0% and 23.4% absolute success-rate gains on real-world single-arm Franka and dual-arm ARX robot platforms.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
action representation
visual dynamics
robot policy learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Dynamics
Action Latent Space
Vision-Language-Action Models
Action VAE
Robot Policy Learning
🔎 Similar Papers
No similar papers found.
Y
Yuan Xu
CASIA, UCAS
Y
Yixiang Chen
CASIA, UCAS
Q
Qisen Ma
CASIA, UCAS
J
Jiabing Yang
CASIA, UCAS
Peiyan Li
Peiyan Li
Ludwig-Maximilians-Universität München
data mininggraph mining
K
Kai Wang
CASIA, UCAS
J
Jianhua Yang
CASIA, UCAS
Jianlou Si
Jianlou Si
alibaba-inc.com
MLLMGenAIAGIEmbodied AI
J
Jun Huang
Alibaba Group
J
Jing Liu
CASIA, UCAS, FiveAges
N
Nianfeng Liu
CASIA, UCAS, FiveAges
Y
Yan Huang
CASIA, UCAS, FiveAges
L
Liang Wang
CASIA, UCAS