🤖 AI Summary
This study addresses the challenge that a single Vision-Language-Action (VLA) model struggles to adapt to varying task states and environments. To overcome this limitation, this work proposes the SWAP framework, which formulates multi-policy dynamic routing as an offline reinforcement learning problem. Specifically, SWAP integrates multiple VLA models via a routing critic network, enabling stepwise online optimal action selection rather than the rigid execution of a fixed policy. Experimental results demonstrate that the proposed approach increases the success rate by 33% on real-world tasks while reducing the effective action sequence length by 28.3%, highlighting its efficacy in enhancing both robustness and efficiency for robotic manipulation.
📝 Abstract
Robot manipulation systems using Vision-Language-Action (VLA) model backbones typically use just one VLA for task execution. However, individual VLAs do not perform well across different task states and environments. We introduce a framework for dynamically composing multiple VLA policies during execution: StepWise Action Policy Routing (SWAP). SWAP formulates policy routing as an offline reinforcement learning problem, learning a routing critic that selects the most appropriate policy at each decision step given the current observation. SWAP enables robots to select new policies to execute online rather than committing to a single policy for the duration of an episode. We evaluate SWAP on both real-world DROID manipulation tasks and LIBERO simulation experiments. SWAP improves over fixed-policy execution and routing baselines, giving absolute improvements in real-world task success up to 33% while reducing successful trajectory robot action step length by 28.3%.