GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the absence of a unified, measurement-based evaluation framework for assessing how accurately existing physics simulators and video world models reproduce real-world physical dynamics. To this end, we introduce GAUGE—a diagnostic benchmark comprising 22 controlled tasks that, for the first time, integrates real-world trajectories, uncertainty annotations, and task-specific observables to enable fine-grained, interpretable evaluation of core physical processes such as collisions, friction, and deformation. Through analyses of generalized trajectory error, consistency with physical laws, and temporal parameter stability, GAUGE reveals significant discrepancies in mainstream simulators during impulsive contacts, rapid cloth motion, and volumetric deformation. Moreover, while video world models can often fit trajectory shapes, they frequently misestimate acceleration, momentum transfer, and oscillation timing.
📝 Abstract
Physics engines facilitate large-scale training and evaluation for embodied intelligence, while generative video world models are emerging as implicit simulators of future states and interactions. However, existing evaluations of physical fidelity are often conducted in isolation and rely heavily on perceptual similarity or human judgments, providing limited insight into which physical principles or parameters are violated. We introduce GAUGE, a real-world-grounded diagnostic benchmark for jointly evaluating how numerical simulators and generative video world models reproduce or deviate from real-world physics. It comprises 22 controlled task families covering rigid bodies, flexible cables, textiles, and volumetric deformable objects. Grounded in real-world trajectories and paired with calibrated physical metadata, uncertainty annotations, and task-specific observables, these tasks cover fundamental physical processes including collision, friction, momentum transfer, oscillation, self-contact, and deformation across diverse materials and conditions. We benchmark Isaac Sim, Genesis, and Newton on 14 task families using generalized trajectory errors, and evaluate 6 image-to-video models on 5 rigid-body tasks by testing physical-law consistency and the temporal stability of inferred parameters. Our results reveal no uniformly faithful physics engine, with the largest discrepancies arising in impulsive contact, rapid textile motion, and volumetric deformation. We further find that video world models can produce trajectories with the expected equation form while recovering incorrect accelerations, momentum transfer, and oscillation timing. GAUGE lays the groundwork for developing more physically faithful simulators and world models for embodied intelligence.
Problem

Research questions and friction points this paper is trying to address.

physical fidelity
simulation engines
video world models
benchmark
real-world grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

physical fidelity
simulation benchmark
video world models
real-world grounding
physics evaluation
🔎 Similar Papers
2024-09-10arXiv.orgCitations: 0
S
Shuai Wang
Shanghai Artificial Intelligence Laboratory
Y
Yaxin Feng
Shanghai Artificial Intelligence Laboratory, Hong Kong University of Science and Technology
X
Xuekun Jiang
Shanghai Artificial Intelligence Laboratory
S
Shihan Tian
Shanghai Artificial Intelligence Laboratory
N
Ningyu Yan
Shanghai Artificial Intelligence Laboratory, Hong Kong University of Science and Technology
X
Xing Shen
Shanghai Artificial Intelligence Laboratory
Chaoyang Lyu
Chaoyang Lyu
Shanghai AI Laboratory
H
Hui Wang
Shanghai Artificial Intelligence Laboratory, Shanghai Jiao Tong University
Yunsong Zhou
Yunsong Zhou
Shanghai Jiao Tong University
Embodied AIGenerative Models
Hanqing Wang
Hanqing Wang
Shanghai AI Laboratory
Embodied AIComputer VisionRobotics
J
Jiangmiao Pang
Shanghai Artificial Intelligence Laboratory
Yang Xiang
Yang Xiang
Professor of Mathematics, Hong Kong University of Science and Technology
Applied and Computational Mathematics
Xing Gao
Xing Gao
Shanghai Artificial Intelligence Laboratory
Graph LearningRobotic LearningAutonomous Driving
Chunhua Shen
Chunhua Shen
Zhejiang University
Computer VisionMachine Learning
Weinan Zhang
Weinan Zhang
Professor, Shanghai Jiao Tong University
Reinforcement LearningAgentsData Science