Attention to details, logits to truth: visual-aware attention and logits enhancement to mitigate hallucinations in LVLMs

📅 2026-02-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the issue of hallucination in large vision-language models, which often arises from diffuse visual attention that inadvertently amplifies irrelevant visual tokens. Existing mitigation strategies typically enhance all visual token attentions uniformly, thereby introducing extraneous noise. To overcome this limitation, the authors propose a training-free attention intervention method that analyzes submatrices of vision-text cross-attention to identify task-relevant visual tokens and selectively reweights their attention scores. Furthermore, visual attention signals are integrated into the beam search decoding process to reinforce visual consistency in generated text. Experiments demonstrate that this approach significantly suppresses hallucinations across mainstream large vision-language models while preserving the accuracy and fluency of the generated outputs.

Technology Category

Computer Vision: Large Vision ModelsMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language Models

Application Category

Search and Retrieval-Augmented AI: Large language models for searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Existing Large Vision-Language Models (LVLMs) exhibit insufficient visual attention, leading to hallucinations. To alleviate this problem, some previous studies adjust and amplify visual attention. These methods present a limitation that boosting attention for all visual tokens inevitably increases attention to task irrelevant tokens. To tackle this challenge, we propose a training free attentional intervention algorithm to enhance the attention of task-relevant tokens based on the argument that task-relevant tokens generally demonstrate high visual-textual similarities. Specifically, the vision-text cross-attention submatrices, which represent visual-textual correlations, are extracted to construct the reweighting matrices to reallocate attention. Besides, to enhance the contribution of visual tokens, we inject visual attention values into the beam search decoding to identify solutions with higher visual attention. Extensive experiments demonstrate that this method significantly reduces hallucinations across mainstream LVLMs, while preserving the accuracy and coherence of generated content.
Problem

Research questions and friction points this paper is trying to address.

hallucinations
visual attention
Large Vision-Language Models
task-relevant tokens
vision-text alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

visual-aware attention
logits enhancement
hallucination mitigation
cross-attention reweighting
training-free intervention
J
Jingyi Wang
Fujitsu Research and Development Center, Beijing, China
Fei Li
Fei Li
Department of Radiology and Huaxi MR Research Center (HMRRC), West China Hospital, Sichuan Universit
RadiologyPsychiatryMRINeuroimageBrain
R
Rujie Liu
Fujitsu Research and Development Center, Beijing, China