🤖 AI Summary
This study addresses the sample efficiency bottleneck in online reinforcement learning for Vision-Language-Action (VLA) models, which arises from low-quality state representations. To this end, we propose a dynamic routing mechanism operating across tokens and layers. By leveraging learnable routing tokens and lightweight layer routers, our method adaptively aggregates multi-depth visual-language features from a frozen VLA to extract action-relevant information, thereby constructing high-quality state representations. This approach is jointly optimized through expert demonstration initialization and critic feedback. Experimental results demonstrate that the proposed method improves the Area Under the Curve (AUC) by 23.7% in simulation tasks and achieves up to a 108.9% gain in real-world robot experiments, significantly enhancing the sample efficiency of policy learning.
📝 Abstract
Vision-Language-Action (VLA) models provide strong behavioral priors for robotic manipulation, yet efficiently adapting them to downstream tasks remains challenging. Recent work addresses this challenge by adapting frozen VLAs through online reinforcement learning (RL), whose sample efficiency depends on the quality of the state representation used by the actor and critic. Existing methods construct such representations either with VLA-independent visual encoders or through fixed compression of internal VLA representations. Neither design explicitly extracts the task-specific action-relevant VLA features most useful for downstream action refinement and action-value estimation, therefore limiting sample efficiency. To address this limitation, we introduce eRLT, which constructs an effective state representation by routing task-specific action-relevant information across both tokens and layers of the frozen VLA. Specifically, learned routing tokens dynamically aggregate visual-language features at multiple depths, while a lightweight layer router combines these summaries into a fixed-dimensional RL token. The routing module is initialized using expert demonstrations to capture features predictive of expert actions and then refined using critic feedback from online interactions for action-value estimation. Across seven LIBERO and RoboTwin tasks, eRLT improves mean normalized learning-curve AUC by up to 23.7% over representative baselines. Real-robot experiments on USB connector insertion and motherboard ribbon-cable insertion further show AUC improvements of 108.9% and 46.7%, respectively, over the strongest baseline.