Visual Evidence Under Cross-Examination: Evaluating and Controlling Decision-Level Evidence Use in Vision-Language Models

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation in vision-language models where visual evidence influences outputs without being precisely bound to specific candidates. To resolve this, the authors propose RIVET, an interface that controls evidence utility, and introduce CROSS-Bench alongside a novel "cross-examination" evaluation paradigm to define candidate-bound visual contributions. By integrating multi-backbone frozen fine-tuning, the approach effectively decouples task accuracy from evidence ownership. Experimental results demonstrate that the normalized effect transmission rate reaches 0.651, while average accuracy improves by 5.70 percentage points across four frozen backbones, significantly optimizing evidence utilization efficiency.
📝 Abstract
Vision-language models increasingly reason through crops, regions, and tool-produced observations. Yet an observation can influence the answer without benefiting the candidate it supports. We study candidate-bound visual contribution: valid evidence should help, invalidating its supporting relation should remove its additional effect, and valid rebinding should redirect that effect to the newly supported candidate. We introduce CROSS-Bench, a benchmark of 28,000 decision problems, with matched invalidation and rebinding tests on a dedicated evaluation subset. Our RIVET interface preserves evidence identity and uncertainty, composes a candidate-conditioned response, and separately controls its strength. Shared-evidence experiments show that task accuracy and evidence ownership can diverge. Under matched capacity and training, RIVET increases normalized effect transfer from 0.512 to 0.651 where clean evidence has a positive effect. The advantage persists on common evaluation examples and across repeated decision-layer fits. With evidence predicted from raw inputs, RIVET improves CROSS-Bench accuracy by an average of 5.70 pp across four frozen backbones, relative to the same models without auxiliary evidence. These results separate the utility of visual evidence from the candidate-specific destination of its effect.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Visual Evidence
Decision-Level Reasoning
Evidence Attribution
Candidate-bound Contribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Visual Evidence
CROSS-Bench
RIVET
Decision-Level Reasoning
H
Huiyao Zhang
Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences; University of Chinese Academy of Sciences
J
Jin Bai
Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences; University of Chinese Academy of Sciences
Z
Zilong Su
University of Science and Technology of China
R
Rui Guo
Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences; University of Chinese Academy of Sciences
C
Chaofan Qin
Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences; University of Chinese Academy of Sciences
J
Jinze Lv
Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences; University of Chinese Academy of Sciences
W
Wenhui Yu
Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences; University of Chinese Academy of Sciences
Hongfei Wang
Hongfei Wang
Communication University of China
audiodeep learning
Y
Ye Li
Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences; University of Chinese Academy of Sciences