ROT: Rotating Hidden States towards Contextual Vectors for Hallucination Mitigation in LVLMs

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses object hallucination in large vision-language models, noting that existing methods relying on indirect attention weight intervention fail to resolve the issue fundamentally. We propose ROT, a training-free framework that, for the first time, focuses on hidden state vectors following self-attention, revealing that hallucinations stem from semantic deviations away from the multimodal context plane. Specifically, ROT detects semantic deviations in intermediate layers and employs geometry-driven norm-preserving rotations to calibrate anomalous representations onto the local context plane, alongside a smoothing mechanism to stabilize generation trajectories. Experiments demonstrate that ROT effectively suppresses hallucinations across multiple benchmarks for models of varying architectures and scales, providing an efficient and generalizable geometrically grounded generation solution.
📝 Abstract
Large Vision-Language Models (LVLMs) frequently suffer from object hallucination. Existing training-free interventions primarily manipulate attention weights, which indirectly affect the deep semantics reaching the final predictive layers. In this work, we shift our focus to the hidden state vectors extracted after self-attention and residual addition. Empirical analysis reveals that hallucinated tokens do not simply over-rely on linguistic priors; instead, they exhibit an anomalous contextual deviation, showing significantly lower similarities to both textual and visual contexts in intermediate layers. Motivated by this, we propose ROT, a layer-specific, training-free framework. ROT dynamically detects semantic deviation in the middle layers and applies a norm-preserving rotation to steer the hidden states back toward the local multimodal context plane spanned by the contexts. For subsequent layers, a representational smoothing mechanism is introduced to stabilize the calibrated trajectory. Extensive experiments on multiple benchmarks demonstrate that ROT consistently reduces hallucinations across various model architectures and scales, offering an efficient, geometry-driven solution for grounded generation.
Problem

Research questions and friction points this paper is trying to address.

Object Hallucination
Large Vision-Language Models
Contextual Deviation
Hidden States
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hallucination Mitigation
Hidden State Rotation
Training-free Intervention
Large Vision-Language Models
Contextual Deviation
🔎 Similar Papers
2024-10-06Conference on Empirical Methods in Natural Language ProcessingCitations: 33
💼 Related Jobs
No related jobs found.
Y
Yijing Du
Harbin Institute of Technology, Shenzhen
X
Xiangcheng Zhan
Harbin Institute of Technology, Shenzhen
Shuo Yang
Shuo Yang
Professor, Harbin Institute of Technology (Shenzhen)
Data-Centric AITrustworthy AIMachine LearningComputer Vision