π€ AI Summary
This work addresses the high data acquisition cost of dual-arm manipulation, which typically relies on extensive dual-arm demonstration datasets. The authors propose ExS2D, a novel framework that, for the first time, enables reliable bimanual coordination using only a few single-arm demonstrations and without any dual-arm data. By parsing text instructions into temporally constrained subtasks, the method leverages subtask-guided action mapping and a multimodal large language model coordinator to achieve collision-free, synchronized motion planning and task allocation. Its core innovations lie in explicit temporal modeling and a multimodal coordination mechanism. Experiments show a 54.4% reduction in average execution steps in simulation while matching the success rate of single-arm baselines; real-robot evaluations across four tasks further validate its effectiveness under zero dual-arm demonstration conditions.
π Abstract
Dual-arm manipulation can improve throughput via parallel execution, but collecting bimanual demonstrations for training is costly and difficult. We present ExS2D, a hierarchical action expansion framework that enables dual-arm manipulation from single-arm supervision. ExS2D first generates structured subtasks from textual instructions while explicitly capturing temporal precedence. It then grounds each subtask into executable actions through subtask-guided action mapping in observation. Finally, precedence-aware action allocation and synchronized planning are performed by a multimodal large language model driven coordinator to select collision-free dual-arm executions. Simulation experiments demonstrate that ExS2D reduces the average execution steps by 54.4% while maintaining a comparable success rate to a single-arm baseline. Real-robot experiments on four tasks further demonstrate the reliability of ExS2D for dual-arm execution under few-shot single-arm samples, while using zero bimanual demonstrations.