Score
Designs and implements evaluation methods, benchmarks, and quantitative metrics that measure how well a model-generated or observed physical scenario conforms to real-world physical laws, including hierarchical test suites and physics-grounded scoring protocols. Builds analysis tools that produce fine-grained per-attribute plausibility scores and localize or attribute specific violations (e.g., which object, frame, or property breaks a conservation, kinematic, or contact constraint).
Current video generation models often produce physically implausible outputs—visually realistic yet violating fundamental physical laws. To address this, we propose the first “three-tier physical cognition” taxonomy—comprising basic schema perception, passive physical knowledge acquisition, and active world simulation—grounded in cognitive science to systematically characterize the evolution of physical modeling in video generation. Methodologically, our framework integrates diffusion-based architectures, physics-inspired motion representations, knowledge embedding mechanisms, and a multi-level physical consistency evaluation benchmark. Key contributions include: (1) establishing the first systematic review framework explicitly targeting physical cognition in video generation; (2) advancing the field from visual fidelity toward human-like physical understanding; and (3) clarifying the synergistic progression path among interpretability, controllability, and physical consistency—providing a structured roadmap for both academic research and industrial development.
This work addresses the lack of fine-grained, auditable evaluation methods for physical reasoning in current video generation models, which hinders diagnosis of their failures with respect to specific physical laws. The authors propose the first physics-grounded benchmark for video generation, encompassing 13 categories of physical principles—including solid mechanics, fluid dynamics, and optics—and comprising 250 prompts paired with expected outcomes. Leveraging a social science-inspired experimental design, they collect 5,796 annotation sets from 459 human annotators, yielding over 37.4K fine-grained labels. Built upon this data, they introduce PhyJudge-9B, an open-source vision-language model judge that enables interpretable, reproducible evaluation with low bias (reducing bias to 3.3% relative to Gemini-3.1-Pro) and high reliability (Spearman correlation > 0.90).
This work addresses the challenge that existing automatic code generation methods often produce structurally invalid or physically inconsistent models, which are unsuitable for engineering simulation. To ensure physical consistency and simulatability, the authors propose a procedural modeling framework that integrates domain knowledge injection, constraint-guided fine-tuning, and closed-loop simulation validation. Key contributions include CivilInstruct—the first instruction-following dataset tailored for structural engineering—along with a two-stage fine-tuning strategy and MBEval, a validation-driven evaluation benchmark. Experimental results demonstrate that the proposed approach significantly outperforms baseline methods across multiple rigorous metrics, effectively suppressing hallucinations and constraint violations, and enabling the direct use of generated models in structural dynamics simulations.
Current large language models exhibit significantly weaker causal and procedural reasoning capabilities in real-world physical scenarios compared to human experts. Method: We introduce PHYBench, the first comprehensive benchmark for physics-grounded reasoning, covering six domains—including mechanics and electromagnetism—with 500 hierarchically structured problems. We propose Expression Edit Distance (EED) as a fine-grained evaluation metric, enabling quantitative, step-by-step comparison between model-generated reasoning paths and expert solutions. PHYBench ensures validity and reliability through realistic scenario modeling, multi-level difficulty design, and expert-verified annotations. Contribution/Results: Experiments reveal that even state-of-the-art reasoning models substantially underperform human experts on PHYBench. The benchmark dataset and evaluation framework are publicly released to advance research in physics-aware reasoning and causal inference.
Model fidelity—the degree of correspondence between simulation and reality—lacks a formal, axiomatic foundation in digital engineering, resulting in ambiguous evaluation criteria and poor cross-domain comparability. Method: This paper introduces the first rigorous, verifiable theoretical framework for fidelity assessment, grounded in seven foundational axioms encompassing consistency, measurability, scale invariance, and other essential properties; the framework enables formal verification and comparative analysis of fidelity metrics. Empirical validation is conducted via integration into ground-vehicle modeling, demonstrating feasibility and practical guidance within existing evaluation paradigms. Contribution/Results: The work fills a critical theoretical gap in fidelity science and establishes a universal, standards-ready paradigm for fidelity assessment—directly advancing digital twin development, simulation verification and validation (V&V), and model-based systems engineering. It further provides a clear, principled roadmap for future methodological evolution and standardization.
This work addresses the lack of systematic evaluation of physical reasoning in existing video generation models, which typically focus only on output plausibility. To bridge this gap, the authors propose the first physics-based benchmark for video generation, comprising a task-specific dataset, a three-stage evaluation protocol—spanning perception, modeling, and inference—and a hybrid assessment framework that integrates infographic-guided frame-chain prompting, subjective scoring by multimodal large language models, and objective metrics grounded in physical laws. Experiments across eleven state-of-the-art models reveal a significant performance gap relative to a reliable physics simulator (best score: 0.473), uncovering stage-wise bottlenecks from perception to inference and highlighting a persistent simulation-to-reality discrepancy.
This work addresses the absence of a unified, measurement-based evaluation framework for assessing how accurately existing physics simulators and video world models reproduce real-world physical dynamics. To this end, we introduce GAUGE—a diagnostic benchmark comprising 22 controlled tasks that, for the first time, integrates real-world trajectories, uncertainty annotations, and task-specific observables to enable fine-grained, interpretable evaluation of core physical processes such as collisions, friction, and deformation. Through analyses of generalized trajectory error, consistency with physical laws, and temporal parameter stability, GAUGE reveals significant discrepancies in mainstream simulators during impulsive contacts, rapid cloth motion, and volumetric deformation. Moreover, while video world models can often fit trajectory shapes, they frequently misestimate acceleration, momentum transfer, and oscillation timing.
Existing video generation evaluation methods rely on human ratings or ground-truth reference videos, making it difficult to effectively assess the physical consistency of videos generated by world models and resulting in a significant performance gap when transferring from simulation to real-world tasks. This work proposes the first reference-free automatic evaluation method for physical consistency, integrating both relative and absolute assessment strategies. By leveraging DROID-SLAM and SEA-RAFT, the approach constructs a spatiotemporal consistency metric that precisely localizes the timing and spatial location of physical artifacts. Experimental results demonstrate that videos selected using this method improve downstream task success rates by over 8%, substantially narrowing the sim-to-real performance gap.