OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the hallucination problem in omni-modal large models, where reliance on erroneous evidence generates untraceable errors. To this end, this work proposes a training-free, token-level evidence attribution and rectification method. During inference, the approach employs channel-level interventions to generate token-level evidence confessions, integrating token rescoring with structured evidence dependency analysis to precisely localize and correct hallucinations driven by irrelevant or contradictory evidence. Key contributions include the first token-level evidence attribution mechanism and the construction of OmniHalluBench, a comprehensive multimodal hallucination evaluation benchmark. Experimental results demonstrate that the proposed method effectively mitigates hallucinations across diverse heterogeneous modalities and tasks. The source code and benchmark have been made publicly available.
📝 Abstract
Omni-modal large language models (OmniLLMs) unify text, images, audio, and video, yet hallucinate when generation relies on the wrong evidence. Existing inference-time methods can reduce hallucinations, but rarely reveal which evidence sustains a generated commitment. We introduce OmniConfess, a training-free method for mitigating omni-modal hallucinations. It fixes a candidate response and re-scores it at token resolution under controlled channel-wise evidence interventions, producing a structured token-by-channel confession that reveals the response's evidential dependence. OmniConfess uses this confession to preserve grounded content and correct commitments driven by irrelevant or contradictory evidence. To evaluate OmniConfess, we construct OmniHalluBench, a 3,540-example benchmark built from six datasets spanning text, image, audio, and video settings and both judgment and free-form generation. Experiments show that OmniConfess mitigates hallucinations across heterogeneous modality and task settings. Our code and benchmark are publicly available at https://github.com/RongHuiQiang/OmniConfess.
Problem

Research questions and friction points this paper is trying to address.

Omni-modal hallucination
Large language models
Evidential dependence
Hallucination mitigation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Omni-modal Hallucination
Training-free Mitigation
Token Confession
Evidence Intervention
OmniHalluBench
🔎 Similar Papers