What 30,000 Hours of Ego-centric Video Does Not Teach

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that existing world models struggle to simultaneously achieve high-fidelity agent-object interactions and uniform benefits from data scaling. Through a systematic evaluation using 30,000 hours of first-person video, we reveal that agent modeling saturates rapidly while object dynamics remain difficult to improve, demonstrating that merely increasing data scale yields diminishing returns for object fidelity. To overcome this bottleneck, we propose a decoupled supervision paradigm that reallocates model capacity toward object dynamics, optimized through visual conditioning design and out-of-distribution benchmarking. Experimental results demonstrate that our approach significantly enhances object interaction fidelity. Furthermore, these findings successfully transfer to humanoid robot modeling, validating that strategic optimization of training paradigms is superior to naive data accumulation.
📝 Abstract
World models offer a promising alternative to physics-based simulators, yet remain far from practical deployment. We ask how far scaling ego-centric human video takes them, using a dataset of 30,000 hours spanning over 1,000 scene types and 14,000 contributors. Rather than relying on opaque downstream metrics, we directly evaluate agent and object-interaction fidelity on a challenging out-of-distribution benchmark. Increasing training data by 100x improves both, but unevenly: the agent is modeled well, while object fidelity remains far lower and improves slowly. We show that the agent gains need not come from data, and a careful visual conditioning design saturates fidelity with a fraction of it, which lets us measure object fidelity on its own and discover its saturation point. We then introduce a supervision scheme that shifts capacity from scene appearance toward object dynamics, improving object fidelity though a substantial gap remains. Finally, our conclusions transfer to downstream humanoid modeling. Overall, our results suggest that scaling ego-centric data brings agent modeling close to its limit while leaving its effects on the world far behind, and that closing this gap will depend on how models are trained, not only on how much data they see.
Problem

Research questions and friction points this paper is trying to address.

world models
ego-centric video
object-interaction fidelity
agent modeling
data scaling
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Models
Ego-centric Video
Visual Conditioning
Supervision Scheme
Object Dynamics
🔎 Similar Papers
No similar papers found.