🤖 AI Summary
This study addresses the cross-interference problem in existing vision-language-action models caused by the coupling of bimanual states and intentions. To this end, it proposes a symmetric dual-arm expert architecture comprising a shared VLM backbone and decoupled expert towers. A two-stage explicit-implicit intention routing mechanism is designed to disentangle left- and right-arm behaviors. Furthermore, a lightweight cross-modal task progress prediction module and a cross-attention-based temporal fusion technique are introduced to facilitate coordinated scheduling. This work effectively eliminates dual-arm control interference, substantially improving execution success rates while enabling skill generalization between single- and dual-arm settings as well as cross-motion-domain transfer.
📝 Abstract
Vision-Language-Action (VLA) models provide a unified framework for grounding high-level semantic information into low-level robot actions, enabling scalable robotic manipulation across diverse tasks. However, existing VLA models lack explicit mechanisms to disentangle the states and intents of the two arms, leading to unintended cross-arm interference that degrades task execution success. To address this issue, we propose a symmetric Dual-Arm Expert (DAE) architecture built upon a shared Vision-Language Model (VLM) backbone with decoupled, arm-specific expert towers. Expert selection is carried out through a two-stage dual-arm intent routing scheme, in which experts are routed either by explicit language instructions in the first stage or by implicit visual semantics in the second stage. Moreover, we introduce a lightweight task progress prediction module that leverages cross-attention between the pre-chunk temporal features and semantic representations of proprioceptive and visual observations to accurately estimate frame-wise task completion progress. This module facilitates task progress synchronization to support coordinated scheduling for collaborative multi-robot tasks. Experimental results demonstrate the effectiveness of our model in dual-arm intent routing and the disentanglement of cross-arm interference, and further provide preliminary evidence of emergent skill generalization from single- to dual-arm tasks (as well as the reverse), together with cross-arm motion-domain skill transfer.