🤖 AI Summary
Large Vision-Language Models are prone to hallucination due to cross-modal attention imbalance. This work proposes FLASH, a framework that corrects visual information flow via spectral surgery to mitigate this issue. Specifically, we identify two novel hallucination patterns and design a Spectral Vortex Score to localize anomalous visual attention heads, followed by adaptive frequency-domain local attention shaping. The proposed framework effectively alleviates hallucinations without requiring additional training or contrastive decoding, thereby preserving inference efficiency. Comprehensive evaluations demonstrate that FLASH achieves superior overall performance compared to existing state-of-the-art methods.
📝 Abstract
While Large Vision-Language Models (LVLMs) achieve remarkable success, hallucinations remain a significant barrier to their reliable deployment. Recent studies primarily attribute these issues to cross-modal attention imbalances; most solutions therefore focus on reweighting visual tokens or suppressing language priors. However, such approaches often overlook the spectral characteristics of the visual information flow and frequently rely on Contrastive Decoding (CD), which doubles inference time. Instead of following conventional approaches, we identify two distinct hallucination patterns-Perceptual-Semantic Dissociation and Localized Fixation-and propose FLASH (Frequency-Localized Attention SHaping), a training-free and CD-free framework. FLASH utilizes a Spectral Vortex Score to detect vision heads within multi-head attention layers and applies adaptive spectral modulation to rectify the visual information flow during decoding. Empirical results demonstrate that FLASH achieves a superior balance between performance and efficiency compared to SOTA methods.