eRLT: Efficient VLA Reinforcement Learning via Action-Relevant Token Routing

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the sample efficiency bottleneck in online reinforcement learning for Vision-Language-Action (VLA) models, which arises from low-quality state representations. To this end, we propose a dynamic routing mechanism operating across tokens and layers. By leveraging learnable routing tokens and lightweight layer routers, our method adaptively aggregates multi-depth visual-language features from a frozen VLA to extract action-relevant information, thereby constructing high-quality state representations. This approach is jointly optimized through expert demonstration initialization and critic feedback. Experimental results demonstrate that the proposed method improves the Area Under the Curve (AUC) by 23.7% in simulation tasks and achieves up to a 108.9% gain in real-world robot experiments, significantly enhancing the sample efficiency of policy learning.
📝 Abstract
Vision-Language-Action (VLA) models provide strong behavioral priors for robotic manipulation, yet efficiently adapting them to downstream tasks remains challenging. Recent work addresses this challenge by adapting frozen VLAs through online reinforcement learning (RL), whose sample efficiency depends on the quality of the state representation used by the actor and critic. Existing methods construct such representations either with VLA-independent visual encoders or through fixed compression of internal VLA representations. Neither design explicitly extracts the task-specific action-relevant VLA features most useful for downstream action refinement and action-value estimation, therefore limiting sample efficiency. To address this limitation, we introduce eRLT, which constructs an effective state representation by routing task-specific action-relevant information across both tokens and layers of the frozen VLA. Specifically, learned routing tokens dynamically aggregate visual-language features at multiple depths, while a lightweight layer router combines these summaries into a fixed-dimensional RL token. The routing module is initialized using expert demonstrations to capture features predictive of expert actions and then refined using critic feedback from online interactions for action-value estimation. Across seven LIBERO and RoboTwin tasks, eRLT improves mean normalized learning-curve AUC by up to 23.7% over representative baselines. Real-robot experiments on USB connector insertion and motherboard ribbon-cable insertion further show AUC improvements of 108.9% and 46.7%, respectively, over the strongest baseline.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
Reinforcement learning
Sample efficiency
State representation
Robotic manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action models
Reinforcement Learning
Token Routing
State Representation
Sample Efficiency
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Dehao Huang
Dehao Huang
Southern University of Science and Technology
Robot Grasping and Manipulation
Jianbang Liu
Jianbang Liu
Samsung Robotics eXperience.
J
Jianpan Gao
Samsung Robotics eXperience.
C
Chao Tang
Samsung Robotics eXperience.
Z
Zilang Cen
Beijing Zhongguancun Academy, Beijing, China.
Z
Zedong Dan
Sun Yat-sen University, Guangzhou, China.
J
Jiaheng Wang
Samsung Robotics eXperience.
Tingguang Li
Tingguang Li
Tencent Robotics X
Reinforcement LearningRobotics
Y
Yue Wang
Beijing Zhongguancun Academy, Beijing, China.
Hong Zhang
Hong Zhang
School of Cybersecurity and Computer Science, Hebei University
Big DataEdge ComputingInformation SecurityAI