🤖 AI Summary
This study addresses attention anomalies during the decoding phase of vision-language models, where visual evidence is suppressed, leading to hallucinations. It identifies Prompt-Invariant Sink (PIS) as the core underlying cause. To mitigate this issue, the authors propose a lightweight attention redirection mechanism. By distinguishing between early/late and middle-layer attention behaviors, the method leverages CLIP/ViT backbones to generate token-aligned ROI masks, thereby guiding decoder attention toward query-relevant regions based on PIS characteristics for the first time. This approach significantly enhances the model's visual grounding capabilities while effectively suppressing hallucinations, yielding consistent performance improvements across multiple downstream benchmarks.
📝 Abstract
Recent large vision-language models (VLMs) pair a visual encoder with a large language model (LLM) and perform well on diverse image-text tasks, yet their reliability is often limited by decoder attention pathologies that suppress visual evidence and exacerbate hallucinations. In this paper, we revisit visual attention sinks and uncover a structured, layer-dependent behavior: across prompts, early and late decoder layers exhibit prompt-invariant attention collapse onto the same few image regions, which we term PIS (Prompt-Invariant Sinks), whereas mid layers become prompt-conditioned and drive vision-language alignment. This split suggests that treating sinks as a uniform effect is incomplete. Building on this insight, we propose SAGE (Sink-Aware Guided Emphasis), a lightweight intervention that steers decoder attention away from PIS and toward query-dependent regions of interest (ROIs) using token-aligned ROI masks derived from standard vision backbones such as CLIP, ViT, and DINOv3. Evaluated on diverse vision-encoder + decoder-only LLM VLM families, SAGE improves visual grounding, reduces hallucinations, and yields consistent gains across public downstream vision-language benchmarks, including fine-grained visual discrimination settings where localized evidence is crucial, when instantiated with backbone-derived ROI masks.