🤖 AI Summary
This study addresses the scarcity of demonstration data and the misalignment between human video and action modalities in humanoid robot loco-manipulation. It proposes a three-stage training framework—pre-training, mid-training on robots, and post-training for actions—to transfer visuomotor priors from mixed-modality human data to robot control. Methodologically, this work innovatively integrates unpaired video and motion streams to learn complementary dynamics, and constructs a latent world-action model architecture that fuses frozen visual representations with proprioception via gated cross-attention to generate whole-body coordinated executable actions. Experiments demonstrate that the proposed method achieves the highest overall success rate on the SIMPLE benchmark and matches the best baseline performance on a real-world Unitree G1 humanoid robot.
📝 Abstract
Humanoid loco-manipulation requires coordinated locomotion and manipulation informed by future scene evolution and whole-body motion, yet learning these capabilities is constrained by scarce robot demonstrations. Human video and motion datasets offer scalable supervision, but many contain only video or motion rather than paired video-motion data. Moreover, human motion does not directly specify executable robot actions. We present Being-M0.7, a latent world-action model that transfers visual-motion priors learned from mixed-modality human data to humanoid control through pre-training, robot mid-training, and action post-training. We curate a corpus from more than 10,000 hours of raw human-centric data, integrating video-only, motion-only, and paired video-motion streams to learn complementary visual dynamics and whole-body kinematic structure. Joint prediction of future latent visual states and motion encourages visual representations to encode future kinematics. Robot mid-training adapts this coarse-grained prior to robot viewpoints and body dynamics. During action post-training, an action expert combines visual predictive representations from the frozen, adapted prior with current images and proprioception through gated cross-attention, grounding predictive context in executable whole-body commands. Being-M0.7 achieves the highest aggregate success rate among the compared baselines on SIMPLE and matches the strongest baseline on real-world Unitree G1 loco-manipulation tasks.