PhyProbe: Rethinking Physical Consistency Evaluation in Generated Videos

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited generalization of existing video physical consistency evaluation methods, which often rely on vision-language models or overfit annotated data and struggle to simultaneously achieve reliable relative ranking and absolute scoring. To this end, this work proposes PhyProbe, which employs a frozen pretrained spatiotemporal encoder for feature extraction and maps them to physical violation scores via a lightweight scoring head. Furthermore, it introduces a unified training objective integrating pairwise ranking, noise regression, and anchor calibration, effectively resolving inconsistencies between ordinal rankings and cardinal scores under multi-source heterogeneous supervision. Experimental results demonstrate that PhyProbe outperforms existing methods across most benchmarks, with particularly significant improvements in scenarios lacking explicit correspondences, while exhibiting strong alignment with human judgment.
📝 Abstract
Evaluating the physical consistency of generated videos remains a fundamental challenge. Existing approaches rely on off-the-shelf vision-language models, which can often be myopic to physical dynamics, or fine-tuned evaluators trained on human annotations, which overfit to dataset-specific cues and fail to generalize. A key challenge is that existing supervision sources provide either relative ordering or absolute scores, but not both reliably and consistently across varied settings. To this end, we introduce PhyProbe, an evaluator that extracts features from a frozen pretrained spatio-temporal encoder and maps them to a scalar physical consistency violation score via a lightweight scoring head. PhyProbe is trained through a unified objective combining pairwise ranking, regression on noisy scalar annotations, and anchor-based calibration over a curated set of heterogeneous supervision sources. Experiments show that PhyProbe outperforms prior methods on most pairwise benchmarks spanning real-generated and generated-generated pairs under varying correspondence, with the largest gains in no-correspondence and generated-generated settings where existing fine-tuned evaluators degrade sharply. PhyProbe achieves strong correlation with human judgments, with close agreement between rank-based and linear metrics, indicating that scores are both well ordered and anchored to a stable [0, 1] scale. Further, despite being trained on supervision indicative of physical consistency, without explicit general-preference labels, PhyProbe also performs competitively on human preference benchmarks: consistent with the observation that physics violations are entangled with broader quality degradations.
Problem

Research questions and friction points this paper is trying to address.

Physical Consistency
Generated Videos
Video Evaluation
Vision-Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Physical Consistency Evaluation
Spatio-temporal Encoder
Unified Training Objective
Anchor-based Calibration
Generated Video Assessment
🔎 Similar Papers