PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited ability of current video large language models to reliably assess whether video events adhere to physical laws and the absence of fine-grained, evidence-based evaluation benchmarks. To this end, we introduce PhyCheck—a multi-level, evidence-oriented video question-answering dataset that encompasses three tiers: coarse-grained physical plausibility judgment, fine-grained identification of physical details, and diagnostic reasoning with external causal context. PhyCheck establishes the first systematic evaluation framework for probing models’ understanding of physical principles. Using this benchmark, we fine-tune and evaluate Qwen2.5-VL, finding significant improvements in recognizing surface-level physical consistency but persistent deficiencies in integrating external causal conditions for deeper physical reasoning, thereby revealing a critical limitation in current models’ comprehension of underlying physical mechanisms.
📝 Abstract
Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of physical regularities. However, despite their strong performance on general video understanding tasks, current video-language models still struggle to reliably determine whether an observed event conforms to specific physical laws. Existing benchmarks primarily assess the physical quality of generated videos, providing limited support for systematically evaluating and improving the physical-law understanding of Video Large Language Models (VideoLLMs). To address this gap, we introduce PhyCheck, a video question answering dataset organized at two complementary levels of granularity. The coarse-grained subset asks models to determine whether the phenomenon shown in a video conforms to or violates physical laws, while the fine-grained subset further examines whether models can capture physical details responsible for the violation or compliance. We use these subsets as structured supervision to improve physical understanding. In addition, the dataset contains a diagnostic subset with external causal context that reveal hidden factors affecting physical plausibility, assessing whether models can recalibrate their judgments accordingly. Experiments with Fine-tune Qwen2.5-VL show that training with the proposed data substantially improves the understanding of physical-consistency, while evaluations in the diagnostic subset reveal that current models still have difficulty incorporating additional causal conditions into their decisions. These findings highlight the gap between recognizing surface-level inconsistencies and understanding underlying physical mechanisms, and provide a foundation for evaluating and improving physical understanding in Video-LLMs.
Problem

Research questions and friction points this paper is trying to address.

physical law understanding
video-language models
physical plausibility
embodied intelligence
world models
Innovation

Methods, ideas, or system contributions that make the work stand out.

physical law understanding
fine-grained evidence
video-language models
causal context
structured supervision
🔎 Similar Papers