EgoLAP: Learning from Egocentric Human Data through Language-Action Reasoning

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that first-person human data cannot be directly applied to robot control due to embodiment discrepancies, proposing the EgoLAP framework. Its core innovation lies in replacing low-level actions with cross-embodiment motion intentions and jointly learning human and robot trajectories through a shared language-action chain-of-thought. Furthermore, it introduces a motion-level reasoning mechanism integrating scene geometry, physics, and object affordances. Built upon vision-language-action (VLA) pretraining, this approach constructs a multimodal motion reasoning model. Real-world experiments demonstrate that EgoLAP achieves an average task progress of 80.1%, outperforming alternative action representations by a factor of 2.3 and surpassing composite reasoning formats.
📝 Abstract
Egocentric human data offer a path to scaling robot learning beyond costly robot demonstrations, yet the embodiment gap makes raw human trajectories a poor supervisory target for control. Our key insight is that, although low-level actions are embodiment-specific, their underlying motion intent can capture task-relevant structure that transfers across humans and robots. We introduce EgoLAP, a VLA pre-training framework that jointly learns from human and robot trajectories through a shared language-based action chain-of-thought. EgoLAP expresses motion intent as structured, temporally abstracted language actions and pairs them with motion-level reasoning grounded in scene geometry, physics, and object affordances. Across extensive real-world and simulated experiments, EgoLAP transfers human experience to robot control more effectively than alternative action representations and reaches 80.1% mean real-world task progress, a 2.3x performance gain over alternative action representations. Motion-level reasoning also outperforms a composite reasoning format that combines subtask, object-box, and visual-trace reasoning.
Problem

Research questions and friction points this paper is trying to address.

egocentric human data
embodiment gap
robot learning
motion intent
Innovation

Methods, ideas, or system contributions that make the work stand out.

Egocentric Human Data
Vision-Language-Action (VLA)
Action Chain-of-Thought
Motion-level Reasoning
Embodiment Gap
🔎 Similar Papers
No similar papers found.