🤖 AI Summary
This study addresses the limitation of existing evaluations for vision-language models in causal reasoning, which often conflate genuine reasoning with shortcut learning and consequently overestimate model capabilities. To this end, we construct a constraint-driven visual causal reasoning benchmark incorporating entity symbolization, spatial grounding, and factual adversarial constraints. These mechanisms systematically suppress spurious shortcut cues within an orthogonal framework, enabling multidimensional standardized diagnosis of causal tasks. Experimental results demonstrate that intervention tasks exhibit the most significant performance degradation, while spatial grounding constraints prove the most disruptive. Crucially, our findings confirm that unconstrained performance is not predictive of a model's robustness under constraints. This work establishes a new paradigm for the precise evaluation of multimodal causal reasoning.
📝 Abstract
Vision-language models (VLMs) have demonstrated excellent performance in visual tasks, but their visual causal reasoning capabilities still lack reliable evaluation. Existing evaluations struggle to distinguish whether a model is performing causal reasoning based on visual evidence or relying on statistical correlations for shortcut learning, thereby potentially overestimating their actual capabilities. This paper proposes CCRV-Bench, a constraint-driven visual causal reasoning benchmark for single-image physical scenarios. We construct an orthogonal framework that evaluates four causal task dimensions: causal relation discovery, state prediction, causal diagnosis, and intervention. We further introduce entity symbolization, spatial grounding, the factual adversarial constraint, and minimalist output constraints to reduce shortcut cues while preserving the physical commonsense required by the task. Experiments across 15 multimodal models show that constraint sensitivity is task- and model-dependent: intervention has the largest average effective degradation among the four causal tasks, spatial grounding is the most damaging constraint on average, and the factual adversarial constraint improves DCR for all evaluated models. These results show that unconstrained performance does not determine constrained robustness and that a single aggregate score can obscure distinct failures in causal identification, spatial grounding, and constraint-compliant expression. CCRV-Bench provides a standardized framework for diagnosing image-grounded causal reasoning under controlled constraints. The code is available at https://github.com/0815linyuan/CCRV-Bench-Constraint-Based-Evaluation-of-Causal-Reasoning-in-Vision-Language-Models