Towards a General Humanoid Loco-Manipulation Model via Egocentric Whole-Body Human Data Pretraining

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of whole-body supervision in egocentric videos and the scalability challenges of teleoperation data for humanoid whole-body mobile manipulation. To this end, it introduces the HumanVerse-500 dataset alongside the λ0 policy. Methodologically, a lightweight wearable system is designed to synchronously capture whole-body and hand motions, establishing a shared representation space that decouples human-robot state discrepancies. Leveraging a Vision-Language-Action (VLA) model, human expertise is transferred to robot control through a three-stage pre-training paradigm. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance on the SIMPLE benchmark and four real-world tasks, validating that scaling data significantly enhances generalization capabilities for whole-body control.
📝 Abstract
Humanoid whole-body manipulation has advanced rapidly, enabling policies to coordinate locomotion, posture, bimanual interaction, and dexterous hand movements. Meanwhile, egocentric human videos provide diverse examples of everyday interactions across objects and scenes, offering scalable supervision without robot operation. However, existing supervision from these videos provides limited coverage of whole-body movement and coordination with hand-object interaction, while obtaining such supervision through humanoid teleoperation is also costly and difficult to scale. We therefore explore how human experience can support scalable learning of humanoid loco-manipulation. To support this study, we introduce HumanVerse-500, a 500-hour dataset of diverse human loco-manipulation behaviors in open-world environments, collected with a lightweight wearable system that synchronizes egocentric video with body and hand motion. Building on this dataset, we develop $λ_0$, a whole-body humanoid vision-language-action policy, through three-stage training that first learns interaction from diverse egocentric datasets, then coordinates body and hand motion using HumanVerse-500, and finally adapts the policy to downstream tasks and robot embodiments. Across these stages, $λ_0$ learns a shared representation space for human experience transfer, while domain-specific interfaces handle differences between human and robot states and actions. We evaluate $λ_0$ on SIMPLE and 4 real-world loco-manipulation tasks, achieving state-of-the-art performance, and further analyze its scaling behavior, generalization, and training-stage contributions to understand how human data support downstream whole-body humanoid control. We will release our code, models, and data to support further research.
Problem

Research questions and friction points this paper is trying to address.

humanoid loco-manipulation
whole-body manipulation
egocentric human data
scalable supervision
human experience transfer
Innovation

Methods, ideas, or system contributions that make the work stand out.

Humanoid Loco-Manipulation
Egocentric Human Data
Vision-Language-Action Policy
Shared Representation Space
Three-stage Training
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
C
Chongyang Xu
Alibaba Group; Sichuan University
Zhao Wu
Zhao Wu
NIO
Machine Learning System
J
Jin Chen
Alibaba Group; Shanghai Innovation Institute
Y
Yiming Jiang
Alibaba Group; Beihang University
J
Jinhui Ye
Hong Kong University of Science and Technology
Y
Yuming Jiang
Alibaba Group
Shifeng Zhang
Shifeng Zhang
Institute of Automation, Chinese Academic of Sciences
Computer VisionObject DetectionFace DetectionPedestrian Detection
Z
Ziliang Feng
Sichuan University
Mu Xu
Mu Xu
alibaba
CV LLM VLM VLA RL
Y
Yilun Chen
Alibaba Group
L
Li Lu
Sichuan University
S
Steven C. H. Hoi
Alibaba Group