🤖 AI Summary
This study addresses the reliance on manual effort and poor scalability in constructing interactive simulators from real-world videos for embodied data generation. To this end, we propose an autonomous video-to-simulation benchmark that formulates this task as a software engineering problem for the first time, systematically evaluating the capacity of coding agents to automatically construct and iteratively refine simulation environments through video observation. We incorporate state-of-the-art foundation models and coding agent systems, establishing multidimensional evaluation metrics encompassing geometric and dynamic fidelity. Experiments reveal that even with Claude Opus 5, the task success rate barely exceeds 15%, exposing a significant gap between visual fidelity and functional correctness. These findings underscore the substantial disparity between current automated reconstruction approaches and human-assisted pipelines, highlighting critical challenges in fully autonomous simulator synthesis.
📝 Abstract
Building interactive simulators from real-world observations is a promising way to scale embodied data, but current pipelines still rely heavily on manual environment construction and calibration. We study whether frontier foundation models and coding agents can automate this process end to end. We formulate \emph{autonomous video-to-simulation} as a software engineering task in which an agent observes an embodied video, constructs the corresponding simulated environment and robot behavior, and iteratively refines the result through execution feedback. To evaluate this capability, we introduce \textbf{Video2World}, a benchmark comprising 222 reconstruction instances derived from 189 robot and human demonstration videos. Video2World measures reconstructed worlds along geometric fidelity, dynamic fidelity, and functional correctness, capturing spatial perception, physical reasoning, and executable interaction. Evaluating 9 frontier coding-agent systems reveals a sharp improvement in Task success beginning with Claude Opus 5, rising from below 5\% to over 15\%, while substantial gaps to human-assisted reconstruction remain. We further find that worlds that look better could work worse: better visual fidelity does not always lead to higher task success. This echoes the broader gap between perceptual realism and factual correctness observed in generative models.