Principia: Relational Physics Tests for Video Models

📅 2026-09-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出Principia,通过评估视频中成对物体间的相对运动关系来测试模型的物理推理能力,解决了因帧率、物体尺度和相机校准等因素导致的绝对运动测量难题。
📝 Abstract
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate, object scale, and camera calibration, all of which are often ambiguous or unavailable in generated video. We propose a different approach. When two objects in the same scene obey the same physical law, their motions must satisfy predictable relationships, and these relationships hold independent of calibration. We introduce Principia, a benchmark that evaluates Newtonian physics through relational consistency between paired objects. Principia spans eight phenomena - gravity, restitution, friction, rotational inertia, projectile motion, momentum, pendulum, and mass-spring oscillation - across translational, rotational, collisional, and oscillatory dynamics, using real-world scenes recorded under controlled protocols. We also introduce a calibration-independent consistency score that quantifies physical violation directly in image space. Across thousands of generations from six state-of-the-art video generators, no model exceeds 0.42 on Principia despite all scoring around 0.8 on VBench. Vision-language models are evaluated on their ability to detect relational physics violations, with the best model achieving only 67% accuracy and most performing near chance level.
Problem

Research questions and friction points this paper is trying to address.

physical reasoning
video models
calibration
motion measurement
Innovation

Methods, ideas, or system contributions that make the work stand out.

relational consistency
calibration-independent
physical reasoning
video models
🔎 Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30