🤖 AI Summary
This study addresses the limitation of current video generation models, which often produce visually realistic yet physically implausible content, while existing benchmarks struggle to evaluate latent physical properties such as weight and viscosity. To this end, we introduce PhysicsLENS, a robotics-scenario-based dataset that pioneers a paired comparative evaluation paradigm targeting hidden physical attributes, systematically assessing model comprehension across seven physical domains. Leveraging publicly available robotic video sources, we conduct multi-domain human annotation and cross-model comparisons. Experimental results reveal that 34 out of 47 test cases neglect the specified physical properties, demonstrating that reliance on textual prompts alone fails to substantially improve the physical plausibility of generated content. These findings expose critical blind spots in the physical reasoning capabilities of contemporary video generation models.
📝 Abstract
Reliable video world models could provide scalable predictive environments for robot learning, planning, and evaluation. However, generated robot videos can violate physical principles and complete tasks through physically implausible behavior, limiting their reliability for robot learning and planning. Current video-generation benchmarks exclude physics that are inherently hidden by visuals (e.g., weight, viscosity, friction). Due to this, video models are evaluated on the fidelity of physics, not the underlying accuracy of physics. We introduce PhysicsLENS, a dataset and benchmark for evaluating plausibility of physical properties grounded in robotics. PhysicsLENS uses matched scenario pairs that hold the same conditioning frame and task, while varying underlying physics in the scene description. Scenarios are curated from public robot video sources and annotated across seven physical domains: collision, gravity, momentum, friction, deformation, fluid, and causality. We evaluate across four video generation models, producing over 400 human-annotated labels. Results show that plausible-looking videos often ignore the stated property (34 of 47), and that stating the property lowers plausibility only slightly and not significantly.