DexRoam: Learning Mobile Bimanual Dexterous Manipulation from Egocentric Whole-Body Human Demonstrations

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of robot demonstrations for mobile bimanual dexterous manipulation, where existing methods lose fine-grained coupled structures by oversimplifying human motion. To overcome this bottleneck, we propose a learning system based on first-person whole-body human demonstrations. Specifically, we design a tracking-free VR data collection scheme and an explicit three-stage alignment mechanism encompassing embodiment, action semantics, and temporal dimensions. Integrated with Vision-Language-Action (VLA) models, this approach enables joint human-robot learning while preserving the continuity of whole-body motion. Empirically, our method achieves performance parity using only half the robot demonstrations. Furthermore, it improves the success rates of GR00T and pi0.5 to 56% and 57%, respectively, validating the effectiveness of the proposed alignment strategy.
📝 Abstract
Mobile bimanual dexterous manipulation requires continuous coordination of locomotion, whole-body motion, and finger-level dexterity within a single trajectory, creating a severe robot demonstration bottleneck. Egocentric human demonstrations offer a scalable alternative, but prior approaches ease the transfer by simplifying human motion, discarding exactly the fine-grained, coupled structure such tasks depend on. We present DexRoam, a complete system for learning mobile bimanual dexterous manipulation from human demonstrations, in which whole-body motion remains continuous and coupled throughout the human-to-robot transfer process. To enable scalable collection of whole-body human manipulation demonstrations, we develop a tracker-free capture system using only a consumer VR headset and a head-mounted stereo camera, without external cameras or motion trackers. We then perform three explicit alignment stages---embodiment, action-semantic, and temporal---to map captured motion into the robot action space, preserving fine-grained whole-body motion and allowing human and robot demonstrations to be jointly learned by standard VLA policies. Real-world experiments with different VLA backbones show that human demonstrations consistently improve policy learning across training paradigms, raising average success from 29% to 56% on GR00T N1.7 and from 32% to 57% on pi0.5, while matching robot-only training with half the robot demonstrations. Ablations confirm that each alignment stage is necessary. These results highlight the potential of human demonstrations for scalable whole-body mobile manipulation with preserved fine-grained motion structure.
Problem

Research questions and friction points this paper is trying to address.

mobile bimanual dexterous manipulation
human demonstrations
whole-body motion
demonstration bottleneck
fine-grained motion structure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mobile Bimanual Dexterous Manipulation
Egocentric Human Demonstrations
Whole-Body Motion Alignment
Tracker-Free Capture
Vision-Language-Action Policies
💼 Related Jobs
No related jobs found.