🤖 AI Summary
This study addresses the limitations of Vision-Language-Action (VLA) models, which suffer from poor cross-domain adaptation, weak generalization, and low fine-tuning efficiency due to tri-modal misalignment. To overcome these bottlenecks, this work proposes the AGFT framework, which systematically formalizes the tri-modal alignment problem in VLA models for the first time and establishes its theoretical connection to the optimization compactness of flow matching. Specifically, it constrains multimodal representations through an explicit alignment loss and replaces diffusion models with flow matching to accelerate inference. Extensive evaluations across diverse benchmarks demonstrate that AGFT achieves higher success rates and lower inference latency, validating the critical role of tri-modal alignment in enhancing the robustness of robotic policies.
📝 Abstract
Recent advances in Vision-Language-Action (VLA) models point toward general-purpose robotic intelligence by unifying perception, instruction, and control. Despite impressive progress, existing VLA models often adapt poorly due to \emph{tri-modal misalignment} among vision, language, and action, which weakens action grounding and hurts generalization and fine-tuning efficiency. In this work, we present Alignment-Guided Flow Transformer (AGFT), a novel framework that explicitly enforces tri-modal alignment through a dedicated alignment loss, bridging the representational gap across modalities and enhancing task adaptation. While prior research has predominantly emphasized bi-modal vision--language alignment, we systematically formalize and study tri-modal alignment in VLA models, and provide both ablations and analysis to isolate its role in improving adaptation and robustness. To further accelerate deployment, we adopt a flow-matching objective, enabling substantially fewer inference steps than diffusion-based policies while maintaining accuracy. Theoretically, we establish a quantitative connection between the tri-modal alignment gap and the optimization tightness of flow matching; empirically, experiments on the extensive benchmark show that AGFT achieves superior success rates and lower inference latency compared to SOTA baselines, underscoring tri-modal alignment as a key ingredient for scaling robust VLA manipulation.