RouteRLT: Learning When and Which RL Specialist Should Control a Vision-Language-Action Policy

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决视觉-语言-动作模型在高精度任务中的局限,RouteRLT框架通过学习何时及选择哪个强化学习专家策略来控制通用模型,从而在保持广泛能力的同时提高特定任务的精确度。
📝 Abstract
Vision-language-action (VLA) models provide broad manipulation competence, but often struggle during the precision-critical stages that dominate contact-rich industrial tasks such as connector insertion and cable management. A common remedy is to refine a pretrained VLA with reinforcement learning (RL), enabling task-specific improvement beyond behavior cloning. However, how to preserve its generalist behavior while deciding when RL refinement is needed and which specialized policy should act remains an open question. In this work, we present RouteRLT, a routing framework that learns when and which RL specialist, an RL policy trained for a single precision-critical phase, should take control from a generalist VLA. A phase selector identifies the active controller, a stabilizer suppresses transient switches, and an action-boundary manager handles transitions between chunked policy outputs. We evaluate RouteRLT on multi-object pick-and-place tasks in LIBERO, as well as on a real-world cable pickup and port-insertion task with multiple precision-critical stages. In simulation, the learned routing improves over the base VLA and matches routing with privileged phase boundaries, without accessing those boundaries at deployment. The real-robot evaluation validates automatic routing to both the pickup and insertion specialists under an operator-aligned handoff protocol. Altogether, these results show that learned routing applies RL specialist control where precise adaptation is most valuable while preserving generalist VLA behavior, including recovery from failed execution attempts.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
Reinforcement Learning
Precision-Critical Stages
Generalist Behavior
Specialized Policy
Innovation

Methods, ideas, or system contributions that make the work stand out.

RouteRLT
Reinforcement Learning
Vision-Language-Action
Precision-Critical Tasks
Automatic Routing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Chongyu Zhu
Department of Mechanical and Industrial Engineering, University of Toronto, Toronto, ON, Canada
J
Jaden Hinds
Department of Mechanical and Industrial Engineering, University of Toronto, Toronto, ON, Canada
H
Hyegang Kim
Department of Mechanical and Industrial Engineering, University of Toronto, Toronto, ON, Canada
Juan Sebastian Rojas
Juan Sebastian Rojas
PhD Student, University of Toronto
Reinforcement LearningMachine LearningRoboticsRisk
R
Ramy Elmallah
Department of Mechanical and Industrial Engineering, University of Toronto, Toronto, ON, Canada
Chi-Guhn Lee
Chi-Guhn Lee
University of Toronto
Operations ResearchMarkov Decision ProcessesReinforcement Learning