🤖 AI Summary
This study addresses the hallucination problem in Large Vision-Language Models (LVLMs) and the limitations of existing vision-augmented or contrastive decoding methods by proposing AIMS, a training-free framework. By revealing the intrinsic visual attention tendencies of LVLMs, this method dynamically computes attention head-level steering weights through multi-source contextual affinity, adaptively coordinating visual, prefill textual, and generated contexts to suppress hallucinations. Experimental results demonstrate that AIMS effectively mitigates object hallucinations across diverse LVLM architectures and decoding strategies while maintaining competitive general multimodal capabilities.
📝 Abstract
Hallucination remains a significant challenge in Large Vision-Language Models (LVLMs). Existing training-free methods generally mitigate hallucinations through contrastive decoding or visual enhancement, often increasing the relative influence of visual evidence during generation. This raises a fundamental question: Can LVLMs dynamically regulate the contributions of different context sources to suppress hallucinations? In this work, we investigate and quantify how LVLMs coordinate multiple context sources during decoding and examine how this intrinsic behavior can guide hallucination mitigation. We find that LVLMs exhibit an intrinsic vision-attending tendency that can guide adaptive visual steering, while textual contexts can also contribute to hallucination mitigation. Motivated by these findings, we propose AIMS (Adaptive Information Multi-source Steering), a lightweight training-free framework that adaptively coordinates visual, prefilled textual, and generated contexts during decoding. Specifically, AIMS constructs compact prototypes for the three context domains and estimates their affinities with the current query to determine head-wise steering weights. The resulting multi-source steering direction is applied to the query representation, enabling adaptive context integration without additional model training or auxiliary forward passes. Extensive experiments across multiple LVLMs and decoding strategies demonstrate that AIMS effectively mitigates object hallucination while maintaining competitive general-purpose multimodal capabilities.