SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing benchmarks struggle to evaluate the arbitration capabilities of multimodal large language models when confronted with conflicting textual and visual evidence. This work introduces SIGNPOST-Bench, a controlled dataset comprising five types of counterfactual images—original, blank, similar, random, and adversarial—and proposes geographic localization error as a continuous diagnostic metric to systematically quantify how models arbitrate conflicting multimodal cues. Leveraging synthetic scene-text interventions, counterfactual image generation, and pairwise distance measurements, the study evaluates 20 models across 25,555 images. Results reveal that adversarial text increases median localization error by 4.8× and biases 6.5–20.1% of predictions toward the injected target, demonstrating that performance on clean inputs does not predict adversarial robustness and consistently confirming a strong text-guided bias in model behavior.
📝 Abstract
Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a controlled counterfactual benchmark for evaluating text-vision conflict resolution. Each source image is transformed into a counterfactual quintuplet of Original, Blank, Similar, Random, and Adversarial variants. Synthetic, localized scene-text interventions are designed to preserve non-textual content, enabling paired measurements of changes in localization performance and directed shifts toward geographic targets introduced by conflicting text. SIGNPOST-Bench contains 5,111 counterfactual groups and 25,555 image variants from four datasets. We evaluate 20 MLLMs from seven providers. Compared with Original images, Adversarial variants raise median localization error from 282 km to 1,347 km, a 4.8-fold increase. Among geocodable adversarial samples, 6.5-20.1% of predictions lie less than 50 km from the injected target across models, and every evaluated model exhibits a positive mean paired reduction in target distance from Blank to Adversarial. Compatible, unrelated, and conflicting text replacements produce distinct effects on model predictions, while clean-input localization performance does not fully predict robustness to conflicting text. These results establish visual geolocation as a continuous diagnostic of scene-text arbitration and provide a controlled framework for evaluating how MLLMs resolve conflicting multimodal evidence.
Problem

Research questions and friction points this paper is trying to address.

multimodal large language models
text-vision conflict
conflict resolution
benchmarking
scene-text arbitration
Innovation

Methods, ideas, or system contributions that make the work stand out.

counterfactual benchmark
text-vision conflict
multimodal arbitration
scene-text intervention
visual geolocation