Visual Credit Audit for Multimodal Spatial Reasoning

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of existing spatial reasoning benchmarks, which often fail to verify whether models genuinely rely on relational visual evidence, leading to correct answers lacking perceptual grounding. The authors propose a training- and label-free Visual Credit Audit (VCA) framework that decomposes multimodal spatial reasoning into three dimensions: correctness, image gain, and relational consistency. Through ablation strategies—including blank/text-only controls, image shuffling, pixel-level relational contrast, and a 3×3 evidence-source factorial design—they introduce Dependency-based Credit Correctness (D-CC) to quantify visual contribution. Experiments across four mainstream multimodal large language models and two benchmarks reveal that 12.73–26.25% of correct predictions lack image support; image shuffling reduces D-CC by 21.25–47.80 points; and under relational inversion, model responses remain 100% consistent, exposing their fragile dependence on visual relationships.
📝 Abstract
Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixed forced-choice interface, Visual Credit Audit (VCA) separates two estimands: whether the benchmark image gives the model's declared decision more support than text-only and blank controls, and whether the model responds to relation-specific visual evidence. The first audit is training- and label-free and does not require an answer flip. Applying labels yields dependence-credited correctness (D-CC); on correct items, it equals same-control gold-aligned positive gain, while prediction alignment extends the audit to errors. Across four open MLLMs and two spatial benchmarks, 12.73-26.25% of decisions are correct yet uncredited. Matched same-split image permutation reduces D-CC by 21.25-47.80 points, with every paired 95% interval above zero. Fixed-pixel relation contrasts and a 3x3 evidence-source factorial show why null controls cannot identify relation response. Among controlled correct-but-uncredited agreement decisions, response to relation reversal spans 81.57-100.00%, while 32.11% pooled change answer. Independently audited outcomes on 108 geometry-compatible edits provide a bounded natural-image correspondence check. VCA thereby decomposes benchmark success into correctness, additional image support, and relation-consistent response.
Problem

Research questions and friction points this paper is trying to address.

multimodal spatial reasoning
visual credit
benchmark evaluation
relation-specific evidence
image support
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Credit Audit
multimodal spatial reasoning
dependence-credited correctness
relation-specific visual evidence
control-based evaluation
🔎 Similar Papers
No similar papers found.
F
Feixiang Liu
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences
Qiang Qiu
Qiang Qiu
Purdue University
Computer VisionPattern RecognitionMachine LearningDeep LearningImage Processing
L
Lanbo Sun
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences
N
Nan Wei
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences
H
Huawei Shen
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences
Xueqi Cheng
Xueqi Cheng
Ph.D. student, Florida State University
Data miningLLMGNNComputational social science