MoGround: Measuring and Mitigating Modality Distraction in Vision-Language Models

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses "modality distraction" in vision-language models (VLMs), a phenomenon wherein irrelevant modalities interfere with reasoning and cause answer flipping. To isolate cross-modal interference, we introduce the first rigorous definition of unimodal answerability and construct a dedicated dataset to quantify this effect. Our analysis reveals that the degree of distraction is negatively correlated with modality grounding strength. Building on this insight, we propose a low-cost mitigation strategy leveraging robust vectors in weight space. Experiments demonstrate that our approach significantly reduces distraction rates by 9% to 51% across seven open-source VLMs, while incurring only a marginal 0.1 percentage point drop in standard task accuracy. These results indicate an effective balance between enhanced robustness against cross-modal interference and minimal performance degradation on conventional benchmarks.
📝 Abstract
We release MoGround, a vision-language dataset spanning four visual domains in which the answer to every question is guaranteed to be available from exactly one modality. This guarantee enables us to measure modality distraction, the failure in which a model answers a question correctly from one modality alone and then flips to a wrong answer once irrelevant content from the other modality is added. Existing probes rarely establish single-modality answerability this way, making it hard to isolate distraction in the first place. Across seven open-source VLMs, we find that modality distraction is not universal but model-dependent. The weaker-grounded modality is the more distracted one (r = +0.86), and distraction scales inversely with grounding strength (r = -0.90). The single-modality guarantee also enables a mitigation method that needs to distinguish between relevant and irrelevant context. Trained on one split of MoGround alone, a weight-space robustness vector reduces distraction on all seven models by 9% to 51%, at a cost of only 0.1 average points of accuracy on standard multimodal tasks.
Problem

Research questions and friction points this paper is trying to address.

Modality Distraction
Vision-Language Models
Multimodal Reasoning
Single-Modality Answerability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Modality Distraction
Vision-Language Models
Single-Modality Guarantee
Weight-Space Robustness Vector
MoGround Benchmark