Do Vision Models Learn Physical Constraints or Rendering Shortcuts? A Counterfactual Benchmark for Grounded Physical Consistency

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the tendency of current image editing models to generate content that violates physical laws, such as missing shadows and incorrect reflections. To this end, we construct a counterfactual synthetic data pipeline based on Mitsuba3 to generate an image benchmark containing physical violations, enabling the diagnosis and localization of physical inconsistencies in edited outputs. Baseline evaluations using vision-language models, including LLaVA and Qwen, reveal that while existing models achieve an F1 score of 67.8% on standard test sets, performance drops sharply to 40.8% under intervention shifts. These findings expose an overfitting phenomenon wherein vision models rely on rendering shortcuts rather than genuinely understanding physical constraints. Ultimately, this work establishes a new paradigm for evaluating and enhancing the physical consistency of generative models.
📝 Abstract
Modern image editing models can satisfy a text instruction while breaking the physics of the edited scene. A new object may cast no shadow, a mirror may fail to reflect visible geometry, or an object may float above a surface that should support it. We study physical plausibility diagnosis, detecting whether an edited image violates scene physics, naming the violation type, localizing the affected region, and explaining the failure in language. We introduce a counterfactual benchmark whose controlled synthetic component uses Mitsuba~3 to generate 5,500 images from 500 scene families. Each family contains one clean image and ten matched violations involving shadows, reflection, support, surface response, and occlusion. The renderer pipeline provides category labels, affected-region masks and boxes, scene metadata, and explanation targets. We use LLaVA-1.5-7B, Qwen2.5-VL-7B, and InternVL3.5-8B as diagnostic baselines rather than proposed methods. On a 1,650-image synthetic test set, the adapted baselines reach 64.0--67.8\% category macro-F1 on standard held-out scenes. For LLaVA-1.5-7B, category macro-F1 falls from 64.0\% on the standard split to 40.8\% under intervention shift. This gap shows that high in-distribution accuracy partly reflects cues tied to rendering and counterfactual construction.
Problem

Research questions and friction points this paper is trying to address.

physical consistency
image editing
counterfactual benchmark
physical plausibility diagnosis
rendering shortcuts
Innovation

Methods, ideas, or system contributions that make the work stand out.

counterfactual benchmark
physical consistency
rendering shortcuts
vision-language models
physically plausible diagnosis
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
M
M. Moein Esfahani
Tri-institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS), Georgia State University, Georgia Institute of Technology, and Emory University, Atlanta, GA, USA
S
Sepehr Salem
Georgia State University, Atlanta, GA, USA
Mohammed Alser
Mohammed Alser
TT Assistant Professor, GSU, ALSER Lab
BioinformaticsMetagenomicsComputational GenomicsComputer Architecture
V
Vince Calhoun
Tri-institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS), Georgia State University, Georgia Institute of Technology, and Emory University, Atlanta, GA, USA