ATI-VLA: Action-Centric Predictive Vision-Language-Action Models via Actionable Alignment Then Adaptive Injection

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance bottlenecks in predictive vision-language-action models caused by modality misalignment and joint optimization conflicts, proposing an action-centric predictive framework. The core innovation lies in introducing a novel two-stage paradigm termed "executable alignment followed by adaptive injection." Specifically, this approach constructs a unified discrete latent space via a shared codebook to align observation and action representations, and subsequently employs a lightweight side network to adaptively inject predicted latent variables as priors for action generation. Extensive evaluations demonstrate that the proposed method achieves state-of-the-art performance across both simulated and real-world robotic tasks while significantly accelerating model convergence.
📝 Abstract
Predictive Vision-Language-Action (VLA) models aim to improve robotic manipulation via future observation or world dynamics forecasting. However, existing approaches often fail to realize this potential and underperform direct action prediction models. We argue that these limitations stem from modality misalignment between observations and actions, together with joint optimization conflicts that drive learning away from an action-centric objective. To this end, we introduce ATI-VLA, an Action-Centric Predictive Vision-Language-Action framework via Actionable Alignment Then Adaptive Injection. Specifically, it follows a two-step design: 1) Actionable Representation Alignment via a Shared Codebook. It aligns predictive observation and action representations by mapping both modalities into a shared discrete latent space via a unified codebook, making predictive observation latents readily usable for action generation and mitigating modality misalignment. 2) Action-Centric Adaptive Injection of Predictive Latents. Building upon this, it then injects predictive observation latents into action decoding as explicit predictive priors via a lightweight adaptive side-path, enabling adaptive predictive guidance under a single action-centric objective. Extensive experiments on both simulation and real-world robotic tasks demonstrate that ATI-VLA achieves state-of-the-art performance with faster convergence.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
robotic manipulation
modality misalignment
predictive modeling
action-centric objective
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Models
Shared Codebook
Representation Alignment
Adaptive Injection
Robotic Manipulation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yijie Zhu
Harbin Institute of Technology, Shenzhen
Rui Shao
Rui Shao
Professor, Harbin Institute of Technology (Shenzhen)
Computer VisionMultimodal LLMEmbodied AI
Jie He
Jie He
Georgia Institute of Technology
Climate Science
W
Wei Li
Harbin Institute of Technology, Shenzhen
B
Bo Zhao
Institute for Artificial Intelligence, Great Bay University
Y
Yelin Wang
Institute for Artificial Intelligence, Great Bay University
X
Xiaochen Yuan
Macao Polytechnic University
Tao Tan
Tao Tan
FCA MPU
Medical Imaging AI
M
Miao Zhang
Harbin Institute of Technology, Shenzhen
Xiaojiang Peng
Xiaojiang Peng
Shenzhen Technology University
Computer VisionFacial Expression RecognitionMultimodal Emotion Recognition
Zitong Yu
Zitong Yu
U.S. Food and Drug Administration
Medical imagingDeep learningMachine learningImage reconstruction