Look Where You Say You're Looking: Self-Grounded Attention for Visual Reasoning

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the misalignment between visual attention and regions referenced in reasoning text within vision-language models (VLMs). To this end, it proposes Self-Saliency, a method that dynamically identifies target regions based on model-generated reasoning content for the first time. By integrating visual grounding techniques with a self-supervised learning mechanism, Self-Saliency performs post-training optimization on VLMs to strengthen the alignment between visual attention and semantic representations. Experimental results demonstrate that the proposed approach achieves the best average ranking and scores across 25 visual reasoning benchmarks. Furthermore, it effectively mitigates the inherent geometric biases of these models, yielding substantial improvements in the reasoning capabilities of VLMs.
📝 Abstract
We introduce Self-Saliency, a method for training Vision-Language Models (VLMs) to increase the alignment between their visual attention and the image regions mentioned in their reasoning. Self-Saliency uses a grounding model to localize the objects mentioned in each reasoning step and treats the resulting areas as supervision for the model's visual attention. Previous work on steering visual attention determines target image regions based solely on the image and question. In contrast, we show that conditioning the target regions on the model's generated reasoning improves downstream performance. For proper evaluation, we build a unified, broad suite of 25 visual reasoning benchmarks, where we reproduce the results of previous methods. We find that Self-Saliency significantly outperforms both prior attention-steering methods and baselines that ground image-level text, achieving both a better average rank and a better mean score. Post-training analysis shows that the model primarily adapts its reasoning text to existing attention patterns, producing shorter steps that refer to larger regions. Nevertheless, when controlling for generated text, attention to grounded regions increases significantly across the relevant layer. Finally, we identify a consistent geometric bias in VLM visual attention toward the image border. However, our ablations show that Self-Saliency's gains cannot be explained by simply aligning attention with the center of the image, highlighting the importance of aligning visual attention with the regions mentioned in the model's reasoning.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Visual Reasoning
Attention Alignment
Self-Grounded Attention
Geometric Bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Saliency
Vision-Language Models
Visual Attention
Grounding Model
Visual Reasoning