🤖 AI Summary
This study addresses the limitation in vision-language models where visual evidence influences outputs without being precisely bound to specific candidates. To resolve this, the authors propose RIVET, an interface that controls evidence utility, and introduce CROSS-Bench alongside a novel "cross-examination" evaluation paradigm to define candidate-bound visual contributions. By integrating multi-backbone frozen fine-tuning, the approach effectively decouples task accuracy from evidence ownership. Experimental results demonstrate that the normalized effect transmission rate reaches 0.651, while average accuracy improves by 5.70 percentage points across four frozen backbones, significantly optimizing evidence utilization efficiency.
📝 Abstract
Vision-language models increasingly reason through crops, regions, and tool-produced observations. Yet an observation can influence the answer without benefiting the candidate it supports. We study candidate-bound visual contribution: valid evidence should help, invalidating its supporting relation should remove its additional effect, and valid rebinding should redirect that effect to the newly supported candidate. We introduce CROSS-Bench, a benchmark of 28,000 decision problems, with matched invalidation and rebinding tests on a dedicated evaluation subset. Our RIVET interface preserves evidence identity and uncertainty, composes a candidate-conditioned response, and separately controls its strength. Shared-evidence experiments show that task accuracy and evidence ownership can diverge. Under matched capacity and training, RIVET increases normalized effect transfer from 0.512 to 0.651 where clean evidence has a positive effect. The advantage persists on common evaluation examples and across repeated decision-layer fits. With evidence predicted from raw inputs, RIVET improves CROSS-Bench accuracy by an average of 5.70 pp across four frozen backbones, relative to the same models without auxiliary evidence. These results separate the utility of visual evidence from the candidate-specific destination of its effect.