DARE to Mitigate Hallucination: Dual-path Auto-Regressive-aware Editing

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the object hallucination problem induced by autoregressive decoding in large vision-language models. To this end, we propose DARE, a hybrid representation editing framework that introduces a novel dual-path autoregressive-aware editing mechanism. Specifically, DARE integrates teacher-forced textual contrast, decoding state transition analysis, and paired-image visual discrepancies to construct a multidimensional hallucination editing vector, enabling precise rectification of model outputs. Experimental results demonstrate that DARE significantly reduces hallucination rates across multiple benchmarks while preserving multimodal perception capabilities and inference efficiency. The source code has been made publicly available.
📝 Abstract
Large vision-language models (LVLMs) have recently achieved remarkable progress across multimodal tasks, yet object hallucination remains a persistent challenge where models generate descriptions inconsistent with the visual input. Recent work mitigates hallucinations through training-free representation editing, typically by constructing hallucination-related directions from teacher-forcing (TF) contrasts between hallucinated and truthful responses. However, LVLMs operate through autoregressive (AR) decoding during generation, raising the question of whether TF-based analysis fully reflects the generation dynamics that lead to hallucinated outputs. In this paper, we analyze the relationship between TF-based editing and AR generation behavior and find that TF-based editing alone may be insufficient to capture both decoding dynamics and multimodal interactions associated with hallucinations. To address this limitation, we propose DARE (Dual-path Auto-Regressive-aware Editing), a hybrid hallucination editing framework that integrates two complementary contrast pathways: textual contrasts and image contrasts, together with autoregressive-aware representation signals. Specifically, DARE constructs hallucination editing directions from (1) TF-based textual contrasts, (2) AR-aware representation transitions during decoding, and (3) controlled visual differences between paired images. Extensive experiments on multiple LVLM hallucination benchmarks demonstrate that DARE consistently reduces object hallucinations while preserving multimodal perception capability and inference efficiency. Our implementation code is available at https://github.com/KU-VGI/DARE.
Problem

Research questions and friction points this paper is trying to address.

Object Hallucination
Large Vision-Language Models
Autoregressive Decoding
Teacher-Forcing
Representation Editing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hallucination Mitigation
Autoregressive-aware Editing
Representation Editing
Large Vision-Language Models
Training-free
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.