🤖 AI Summary
This study addresses motion retargeting distortion and dynamic infeasibility in humanoid mobile manipulation caused by morphological discrepancies, proposing the DexWeave framework. Methodologically, it introduces a novel two-stage coupled retargeting mechanism to ensure interaction consistency, alongside a Transformer policy network integrating anatomically structured tokens with directed masked attention. This architecture supports direct end-to-end reinforcement learning without requiring pretraining or knowledge distillation. The proposed framework significantly enhances retargeting fidelity and manipulation performance while accelerating policy convergence. Furthermore, real-world dexterous mobile manipulation capabilities are successfully validated on the Unitree G1 physical platform.
📝 Abstract
Learning dexterous humanoid loco-manipulation from human demonstrations requires transferring not only human motion, but also the coordinated interaction structure underlying the demonstrated behavior. This is challenging because embodiment differences distort the coupling among body motion, wrist placement, finger articulation, and object interaction, while kinematically accurate references may still be difficult to realize under robot dynamics. We present DexWeave, a unified framework that connects interaction-consistent motion retargeting with anatomy-aware whole-body policy learning. DexWeave first employs a two-stage retargeting procedure that initializes body and hand motions with specialized solvers and subsequently performs coupled refinement over the upper-body interaction chain while preserving lower-body support. The resulting references are tracked by an anatomy-aware Transformer policy that represents anatomical regions as structured tokens and uses directed masked attention to model their dependencies, with object information selectively conditioning the upper-body pathway for dexterous interaction. The policy jointly outputs body and dexterous-hand actions and is trained directly with reinforcement learning, without pretrained tracking policies, teacher-student distillation, or subsequent residual refinement. DexWeave improves retargeting fidelity and interaction consistency while achieving higher manipulation performance and faster policy convergence than MLP baselines. We further deploy the learned policies on a physical Unitree G1 humanoid equipped with Inspire dexterous hands, demonstrating dexterous whole-body loco-manipulation in the real world. See our project page (https://dexweave.github.io) for videos.