Train Together or Merge Later? Unifying VLA Experts via a Shared Action Interface

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance degradation observed when merging independent expert models in multi-task Vision-Language-Action (VLA) policies due to incompatible action interfaces. To this end, it proposes PolicyWeave, a framework that identifies interface conflicts as the primary cause of merging failures and resolves them by preserving shared action interfaces combined with post-training to ensure merge compatibility. Furthermore, the method introduces a vision-language context-based sparse weighting strategy that integrates LoRA fine-tuning, contextual scoring, and reinforcement learning optimization to efficiently consolidate independent VLA experts. Evaluated on the RoboCasa365 benchmark, PolicyWeave achieves a 74.1% success rate, surpassing joint RL baselines. Its practical deployment capability is further validated through long-horizon tasks on real-world robots.
📝 Abstract
Co-training offers a straightforward way to build a multi-task vision-language-action (VLA) policy, but can fall short of the performance achieved by training each task independently. The challenge is to retain these task-specific gains in a multi-task policy without joint post-training. Combining independently trained experts through model merging is a natural approach, yet strong individual experts do not necessarily yield a strong merged policy. We identify one source of this incompatibility: task-specific changes to the action interface, comprising action normalization and the action encoder and decoder. We propose PolicyWeave, combining merge-compatible post-training with context-guided sparse merging. During post-training, all experts retain the common base policy's action interface, while task adaptation is restricted to LoRA updates in the hidden layers of the action model. This makes the experts more compatible with existing model merging methods. However, merging all experts can still introduce interference from unrelated tasks at deployment. PolicyWeave scores each expert's LoRA updates using the initial visual-language context, determines the expert set through leave-one-layer-out ranking stability, and forms a sparse weighted merge of the selected updates that remains fixed for current task. We evaluate PolicyWeave with GR00T N1.5 on 18 RoboCasa365 tasks, using only 10% of the target-task demonstrations for supervised fine-tuning (SFT). Preserving the shared action interface raises the average success rate across four static merging methods from 17.0% to 52.8%. PolicyWeave achieves 64.7% success with these SFT experts and 74.1% after task-specific reinforcement learning (RL), compared with 60.7% for joint RL. Further evaluations on LIBERO-10 and an AgileX Piper arm support the deployment of independently learned skills in long-horizon and real-world manipulation.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action (VLA)
Model Merging
Multi-task Policy
Action Interface Incompatibility
Expert Integration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action (VLA)
Model Merging
PolicyWeave
LoRA
Sparse Merging
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.