EgoAlign: Bridging the Human-Humanoid Gap for Long-Range Loco-Manipulation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of training data for humanoid robots and the scale and state mismatches between first-person human demonstrations and robot control. To this end, it proposes EgoAlign, a framework that employs an execution-feedback-guided data construction pipeline integrating scale alignment, control loop refinement, causal replay, and Vision-Language-Action (VLA) model fine-tuning. This approach transforms human demonstrations into supervision signals compatible with whole-body control, enabling policy training without real-robot data and facilitating zero-shot physical deployment. Experimental results demonstrate that the proposed framework significantly improves hand alignment precision and grasping success rates while supporting long-horizon carrying and navigation tasks, thereby substantially reducing data acquisition costs.
📝 Abstract
Egocentric human demonstrations offer an accessible source of task experience, but differences in body scale and controller response, together with missing robot states, limit their value as humanoid training supervision. We present EgoAlign, a data-construction framework that converts these demonstrations into action and state supervision compatible with a general-purpose, continuous whole-body controller, without collecting physical-robot demonstrations. Using the target-robot model and simulator, EgoAlign guides demonstration collection through execution feedback. It preserves locomotion references for visually guided periodic stepping while adapting upper-body interaction geometry through scale alignment and controller-in-the-loop refinement. A final causal replay reconstructs the corresponding robot states and motion-token labels for training with the human observations. We assess the resulting supervision by fine-tuning a vision--language--action model solely on adapted human demonstrations and deploying it zero-shot on a physical humanoid. The resulting policies perform long-range object relocation, navigation to unseen goal positions, and independently evaluated foot interaction. Refinement improves simulated hand alignment and physical pickup success over kinematic alignment alone, while human collection reduces on-site acquisition time relative to teleoperation. https://lambdahumanoid.github.io/EgoAlign/
Problem

Research questions and friction points this paper is trying to address.

Humanoid robot
Egocentric demonstration
Loco-manipulation
Cross-embodiment alignment
Imitation learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

EgoAlign
Humanoid
Loco-Manipulation
Vision-Language-Action Model
Controller-in-the-loop Refinement
🔎 Similar Papers
No similar papers found.
Y
Yiming Jiang
Beihang University
C
Chen Jin
Shanghai Innovation Institute
C
Chongyang Xu
Sichuan University
Y
Yilun Chen
Alibaba Group
A
Aimin Hao
Beihang University
Yisheng He
Yisheng He
HKUST
Computer VisionDeep LearningEmbodied AI