Ego4WAM: What Matters When Scaling Egocentric Human Data for Robot Learning?

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses which attributes of first-person human videos drive learning gains in robotics and how to effectively leverage them. To this end, it proposes a unified world-action model framework that systematically disentangles the effects of data alignment, duration, diversity, supervision signals, and utilization strategies, with validation through real-world robot experiments and closed-loop evaluation in RoboDojo. The findings reveal that data value is determined by the synergy of multiple factors rather than scale alone, and confirm the effectiveness of video supervision without action labels. Furthermore, aligned demonstrations are shown to significantly enhance generalization while reducing the demand for target-domain data, offering empirical guidance for the efficient utilization of human video data in robot learning.
📝 Abstract
Egocentric human data provides a scalable source of experience for robot learning, but varies substantially in human-robot alignment, behavioral coverage, and available supervision. Existing work shows favorable scaling with increasing human data, but it remains unclear which data properties drive downstream robot gains and how to use such data throughout the training pipeline. We present a systematic study of egocentric human data with different alignment and supervision under a unified world-action model framework. With the model backbone fixed, we disentangle the effects of human-robot alignment, data duration and task diversity, action supervision, and data usage strategies. We find that aligned human demonstrations substantially improve out-of-distribution generalization and reduce target-task robot data requirements; data duration and task diversity affect downstream capabilities differently; and video-only supervision remains effective without action labels, providing a strong foundation for subsequent video-action training. We validate these findings through closed-loop policy evaluation on both real robots and RoboDojo. Rather than treating data duration as the sole scaling axis, Ego4WAM shows how alignment, task diversity, available supervision, and usage strategy jointly shape the value of egocentric human data for robot learning.
Problem

Research questions and friction points this paper is trying to address.

egocentric human data
robot learning
human-robot alignment
data scaling
action supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

egocentric human data
world-action model
human-robot alignment
video-only supervision
robot learning
🔎 Similar Papers
No similar papers found.