🤖 AI Summary
This study addresses the challenge of transferring whole-body coordinated mobile manipulation skills to humanoid robots directly from human first-person perspective data. To this end, we propose the first vision-centric skill transfer framework. Methodologically, a coarse-to-fine motion alignment mechanism is designed, integrating kinematic correction with dynamic optimization to enhance end-effector precision and whole-body coordination. Furthermore, visual-language-action models, robotic arm rendering, and training-time image augmentation are incorporated to bridge the visual embodiment gap. Experimental results demonstrate that the proposed framework achieves zero-shot skill transfer across four real-world tasks, yielding performance comparable to teleoperation-based policies while substantially reducing data acquisition costs.
📝 Abstract
Human demonstrations capture diverse scenes and rich whole-body skills without requiring robot teleoperation. Prior work on egocentric transfer has emphasized scene generalization in loco-manipulation under decoupled control, leaving direct transfer of coordinated whole-body skills less explored. We present EgoHumanoid-V2, the first egocentric human-to-humanoid skill transfer framework for coordinated whole-body loco-manipulation. At its core, coarse-to-fine action alignment combines kinematic reference correction with dynamics-aware refinement. It improves end-effector pose accuracy while preserving whole-body coordination. We also use robot-arm rendering and training-time image augmentation to reduce the visual embodiment gap and improve viewpoint robustness. On four real-world tasks, vision-language-action (VLA) policies trained on aligned human data show zero-shot skill transfer without target-task robot demonstrations. Task scores are comparable to those of policies trained on teleoperation data at a lower collection cost. These results support human data as direct skill supervision.