🤖 AI Summary
This study addresses the challenge that existing world models struggle to simultaneously achieve high-fidelity agent-object interactions and uniform benefits from data scaling. Through a systematic evaluation using 30,000 hours of first-person video, we reveal that agent modeling saturates rapidly while object dynamics remain difficult to improve, demonstrating that merely increasing data scale yields diminishing returns for object fidelity. To overcome this bottleneck, we propose a decoupled supervision paradigm that reallocates model capacity toward object dynamics, optimized through visual conditioning design and out-of-distribution benchmarking. Experimental results demonstrate that our approach significantly enhances object interaction fidelity. Furthermore, these findings successfully transfer to humanoid robot modeling, validating that strategic optimization of training paradigms is superior to naive data accumulation.
📝 Abstract
World models offer a promising alternative to physics-based simulators, yet remain far from practical deployment. We ask how far scaling ego-centric human video takes them, using a dataset of 30,000 hours spanning over 1,000 scene types and 14,000 contributors. Rather than relying on opaque downstream metrics, we directly evaluate agent and object-interaction fidelity on a challenging out-of-distribution benchmark. Increasing training data by 100x improves both, but unevenly: the agent is modeled well, while object fidelity remains far lower and improves slowly. We show that the agent gains need not come from data, and a careful visual conditioning design saturates fidelity with a fraction of it, which lets us measure object fidelity on its own and discover its saturation point. We then introduce a supervision scheme that shifts capacity from scene appearance toward object dynamics, improving object fidelity though a substantial gap remains. Finally, our conclusions transfer to downstream humanoid modeling. Overall, our results suggest that scaling ego-centric data brings agent modeling close to its limit while leaving its effects on the world far behind, and that closing this gap will depend on how models are trained, not only on how much data they see.