Mind What Matters for Reasoning: Aligning Cross-Modal Attention via Selective Probability Mass Concentration

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Multimodal large language models are susceptible to hallucinations induced by linguistic priors, resulting in the omission of critical evidence during visual reasoning. This work proposes sPMC, a framework that enhances implicit visual grounding through selective regularization to improve reasoning performance. Its core innovation lies in an adaptive head selection mechanism that constrains only a small subset of attention heads sensitive to visual localization, thereby avoiding direct interference with the reasoning process. Furthermore, by integrating segmentation spatial priors with probability mass concentration techniques, cross-modal attention is optimized as a spatial distribution. Experimental results demonstrate that adjusting merely 3%–15% of attention heads yields average zero-shot performance gains of 3%, reaching up to 11.3%, across multiple benchmarks.
📝 Abstract
Multimodal large language models (MLLMs) achieve strong performance on visual reasoning tasks, yet remain prone to hallucinations and over-reliance on language priors, often generating answers without adequately using task-relevant visual evidence. Existing approaches primarily improve reasoning through reasoning-oriented supervision or inference-time strategies. In this work, we study a complementary question: can multimodal reasoning be improved by strengthening implicit visual grounding without directly supervising the reasoning process? Motivated by the functional specialization of attention heads, we investigate whether reasoning can be improved by guiding only the heads most responsive to visual evidence grounding. We propose Selective Probability Mass Concentration (sPMC), a training framework that identifies grounding-responsive heads and selectively regularizes their text-to-image attention. sPMC treats normalized attention over visual tokens as a spatial probability distribution and encourages the probability mass to be assigned to semantically relevant regions using segmentation-derived spatial priors. Adaptive Head Selection restricts this guidance to visually responsive heads while leaving the remaining heads unconstrained to preserve their complementary functions. Across 6 multimodal benchmark suites, sPMC achieves an average zero-shot improvement of 3% and gains of up to 11.3% across multiple MLLMs while regularizing only 3%-15% of their attention heads. These results demonstrate that targeted guidance of sparse and implicit visual evidence pathways can directly improve multimodal reasoning.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Visual Reasoning
Hallucination
Language Priors
Visual Grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Selective Probability Mass Concentration
Multimodal Reasoning
Cross-Modal Attention
Visual Grounding
Adaptive Head Selection
🔎 Similar Papers
No similar papers found.