🤖 AI Summary
This study addresses the absence of an explicit transition between perception and action representations in Vision-Language-Action (VLA) models by proposing a shared intermediate representation framework. Specifically, this work introduces a masked action autoencoder to extract latent action tokens and designs a bridging module that explicitly aligns visual features with these tokens. This alignment provides a structured starting point for action generation, thereby establishing an interpretable perception-to-action mapping mechanism. Experimental evaluations on the LIBERO benchmark and real-world robotic tasks demonstrate that the proposed architecture significantly improves policy success rates. Notably, the success rate of OpenVLA-OFT increases from 51.4% to 65.0%, validating both the effectiveness and generalization capability of the proposed method.
📝 Abstract
Vision-language-action (VLA) models map visual observations and language instructions to continuous robot actions. This task requires a transition from representations that describe the scene and instruction to representations that support action generation. Many continuous-action VLAs leave this transition implicit and supervise it mainly through the final action-prediction loss. We introduce PAIR, a framework that learns a shared perception-action representation between these two spaces. During training, a Masked Action Autoencoder encodes expert action chunks into horizon-aligned Action Latent Tokens. A Bridge Module extracts task-relevant features from the current visual-language representations. PAIR aligns these features with the Action Latent Tokens to form Bridge Tokens that preserve task information and capture the structure of expert actions. The Bridge Tokens are then projected into the action-token space and injected into the initial Action Tokens, providing an action-ready starting point for Action Expert refinement. At inference, the autoencoder is removed, and the Bridge Tokens are generated only from the current observation and instruction. Experiments on LIBERO, LIBERO-Plus, and CALVIN ABC-D show gains for the evaluated OpenVLA-OFT and VLA-Adapter models. On LIBERO-Plus, PAIR raises VLA-Adapter's success rate from 59.1% to 64.2%. On CALVIN, it increases VLA-Adapter's average completed sequence length from 4.42 to 4.53. Across seven real-world tasks, PAIR raises OpenVLA-OFT's success rate from 51.4% to 65.0%. Representation analyses show that Bridge Tokens retain task information while making continuous-action information accessible before Action Expert refinement. These results support a shared intermediate representation as a useful interface between perception and action in continuous-action VLAs.