🤖 AI Summary
This study addresses the performance bottlenecks in predictive vision-language-action models caused by modality misalignment and joint optimization conflicts, proposing an action-centric predictive framework. The core innovation lies in introducing a novel two-stage paradigm termed "executable alignment followed by adaptive injection." Specifically, this approach constructs a unified discrete latent space via a shared codebook to align observation and action representations, and subsequently employs a lightweight side network to adaptively inject predicted latent variables as priors for action generation. Extensive evaluations demonstrate that the proposed method achieves state-of-the-art performance across both simulated and real-world robotic tasks while significantly accelerating model convergence.
📝 Abstract
Predictive Vision-Language-Action (VLA) models aim to improve robotic manipulation via future observation or world dynamics forecasting. However, existing approaches often fail to realize this potential and underperform direct action prediction models. We argue that these limitations stem from modality misalignment between observations and actions, together with joint optimization conflicts that drive learning away from an action-centric objective. To this end, we introduce ATI-VLA, an Action-Centric Predictive Vision-Language-Action framework via Actionable Alignment Then Adaptive Injection. Specifically, it follows a two-step design: 1) Actionable Representation Alignment via a Shared Codebook. It aligns predictive observation and action representations by mapping both modalities into a shared discrete latent space via a unified codebook, making predictive observation latents readily usable for action generation and mitigating modality misalignment. 2) Action-Centric Adaptive Injection of Predictive Latents. Building upon this, it then injects predictive observation latents into action decoding as explicit predictive priors via a lightweight adaptive side-path, enabling adaptive predictive guidance under a single action-centric objective. Extensive experiments on both simulation and real-world robotic tasks demonstrate that ATI-VLA achieves state-of-the-art performance with faster convergence.