evaluate physical plausibility

Designs and implements evaluation methods, benchmarks, and quantitative metrics that measure how well a model-generated or observed physical scenario conforms to real-world physical laws, including hierarchical test suites and physics-grounded scoring protocols. Builds analysis tools that produce fine-grained per-attribute plausibility scores and localize or attribute specific violations (e.g., which object, frame, or property breaks a conservation, kinematic, or contact constraint).

evaluatephysicalplausibility

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.09
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the lack of fine-grained, auditable evaluation methods for physical reasoning in current video generation models, which hinders diagnosis of their failures with respect to specific physical laws. The authors propose the first physics-grounded benchmark for video generation, encompassing 13 categories of physical principles—including solid mechanics, fluid dynamics, and optics—and comprising 250 prompts paired with expected outcomes. Leveraging a social science-inspired experimental design, they collect 5,796 annotation sets from 459 human annotators, yielding over 37.4K fine-grained labels. Built upon this data, they introduce PhyJudge-9B, an open-source vision-language model judge that enables interpretable, reproducible evaluation with low bias (reducing bias to 3.3% relative to Gemini-3.1-Pro) and high reliability (Spearman correlation > 0.90).

benchmarkinggenerative world modelsphysical reasoning

This work addresses the challenge that existing automatic code generation methods often produce structurally invalid or physically inconsistent models, which are unsuitable for engineering simulation. To ensure physical consistency and simulatability, the authors propose a procedural modeling framework that integrates domain knowledge injection, constraint-guided fine-tuning, and closed-loop simulation validation. Key contributions include CivilInstruct—the first instruction-following dataset tailored for structural engineering—along with a two-stage fine-tuning strategy and MBEval, a validation-driven evaluation benchmark. Experimental results demonstrate that the proposed approach significantly outperforms baseline methods across multiple rigorous metrics, effectively suppressing hallucinations and constraint violations, and enabling the direct use of generated models in structural dynamics simulations.

LLM hallucinationphysical consistencyscientific modeling

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Apr 22, 2025
SQ
Shi Qiu
🏛️ Peking University | Beijing Computational Science Research Center

Current large language models exhibit significantly weaker causal and procedural reasoning capabilities in real-world physical scenarios compared to human experts. Method: We introduce PHYBench, the first comprehensive benchmark for physics-grounded reasoning, covering six domains—including mechanics and electromagnetism—with 500 hierarchically structured problems. We propose Expression Edit Distance (EED) as a fine-grained evaluation metric, enabling quantitative, step-by-step comparison between model-generated reasoning paths and expert solutions. PHYBench ensures validity and reliability through realistic scenario modeling, multi-level difficulty design, and expert-verified annotations. Contribution/Results: Experiments reveal that even state-of-the-art reasoning models substantially underperform human experts on PHYBench. The benchmark dataset and evaluation framework are publicly released to advance research in physics-aware reasoning and causal inference.

Assessing model capabilities across multiple physics disciplinesEvaluating LLMs' physical reasoning in real-world scenariosMeasuring performance gaps between AI and human experts

Axioms for Model Fidelity Evaluation

Jul 30, 2025
ET
Evan Taylor
🏛️ Clemson University

Model fidelity—the degree of correspondence between simulation and reality—lacks a formal, axiomatic foundation in digital engineering, resulting in ambiguous evaluation criteria and poor cross-domain comparability. Method: This paper introduces the first rigorous, verifiable theoretical framework for fidelity assessment, grounded in seven foundational axioms encompassing consistency, measurability, scale invariance, and other essential properties; the framework enables formal verification and comparative analysis of fidelity metrics. Empirical validation is conducted via integration into ground-vehicle modeling, demonstrating feasibility and practical guidance within existing evaluation paradigms. Contribution/Results: The work fills a critical theoretical gap in fidelity science and establishes a universal, standards-ready paradigm for fidelity assessment—directly advancing digital twin development, simulation verification and validation (V&V), and model-based systems engineering. It further provides a clear, principled roadmap for future methodological evolution and standardization.

Addressing ambiguity in simulation-reality consistency assessmentDefining rigorous axioms for model fidelity evaluationEstablishing foundations for future fidelity frameworks

Latest Papers

What's happening recently
View more

This work addresses the lack of systematic evaluation of physical reasoning in existing video generation models, which typically focus only on output plausibility. To bridge this gap, the authors propose the first physics-based benchmark for video generation, comprising a task-specific dataset, a three-stage evaluation protocol—spanning perception, modeling, and inference—and a hybrid assessment framework that integrates infographic-guided frame-chain prompting, subjective scoring by multimodal large language models, and objective metrics grounded in physical laws. Experiments across eleven state-of-the-art models reveal a significant performance gap relative to a reliable physics simulator (best score: 0.473), uncovering stage-wise bottlenecks from perception to inference and highlighting a persistent simulation-to-reality discrepancy.

benchmarkingphysical intelligencephysical laws

This work addresses the absence of a unified, measurement-based evaluation framework for assessing how accurately existing physics simulators and video world models reproduce real-world physical dynamics. To this end, we introduce GAUGE—a diagnostic benchmark comprising 22 controlled tasks that, for the first time, integrates real-world trajectories, uncertainty annotations, and task-specific observables to enable fine-grained, interpretable evaluation of core physical processes such as collisions, friction, and deformation. Through analyses of generalized trajectory error, consistency with physical laws, and temporal parameter stability, GAUGE reveals significant discrepancies in mainstream simulators during impulsive contacts, rapid cloth motion, and volumetric deformation. Moreover, while video world models can often fit trajectory shapes, they frequently misestimate acceleration, momentum transfer, and oscillation timing.

benchmarkphysical fidelityreal-world grounding

Existing video generation evaluation methods rely on human ratings or ground-truth reference videos, making it difficult to effectively assess the physical consistency of videos generated by world models and resulting in a significant performance gap when transferring from simulation to real-world tasks. This work proposes the first reference-free automatic evaluation method for physical consistency, integrating both relative and absolute assessment strategies. By leveraging DROID-SLAM and SEA-RAFT, the approach constructs a spatiotemporal consistency metric that precisely localizes the timing and spatial location of physical artifacts. Experimental results demonstrate that videos selected using this method improve downstream task success rates by over 8%, substantially narrowing the sim-to-real performance gap.

physical consistencyreference-free evaluationsimulation-to-reality gap

Hot Scholars

YL

Yiting Lu

University of Science and Technology of China
VLM,Self-evolving Agent,Reasoning Model
ZL

Ziwei Liu

Associate Professor, Nanyang Technological University
Computer VisionMachine LearningComputer Graphics
ZC

Zhaoxi Chen

Ph.D. Student, Nanyang Technological University
Neural renderingGenerative models
SA

Sameer Antani

National Library of Medicine, National Institutes of Health
Medical ImagingMachine LearningArtificial IntelligenceImage Informatics
HZ

Hongyu Zhang

Chongqing University
Software EngineeringMining Software RepositoriesData-driven Software EngineeringSoftware Analytics