🤖 AI Summary
This study addresses the challenge that first-person human data cannot be directly applied to robot control due to embodiment discrepancies, proposing the EgoLAP framework. Its core innovation lies in replacing low-level actions with cross-embodiment motion intentions and jointly learning human and robot trajectories through a shared language-action chain-of-thought. Furthermore, it introduces a motion-level reasoning mechanism integrating scene geometry, physics, and object affordances. Built upon vision-language-action (VLA) pretraining, this approach constructs a multimodal motion reasoning model. Real-world experiments demonstrate that EgoLAP achieves an average task progress of 80.1%, outperforming alternative action representations by a factor of 2.3 and surpassing composite reasoning formats.
📝 Abstract
Egocentric human data offer a path to scaling robot learning beyond costly robot demonstrations, yet the embodiment gap makes raw human trajectories a poor supervisory target for control. Our key insight is that, although low-level actions are embodiment-specific, their underlying motion intent can capture task-relevant structure that transfers across humans and robots. We introduce EgoLAP, a VLA pre-training framework that jointly learns from human and robot trajectories through a shared language-based action chain-of-thought. EgoLAP expresses motion intent as structured, temporally abstracted language actions and pairs them with motion-level reasoning grounded in scene geometry, physics, and object affordances. Across extensive real-world and simulated experiments, EgoLAP transfers human experience to robot control more effectively than alternative action representations and reaches 80.1% mean real-world task progress, a 2.3x performance gain over alternative action representations. Motion-level reasoning also outperforms a composite reasoning format that combines subtask, object-box, and visual-trace reasoning.