🤖 AI Summary
This study addresses the issue of object hallucination in large vision-language models during decoding, which arises from visual uncertainty and lacks principled image-level false positive control in existing methods. To this end, it proposes CORAL, a framework that models visual uncertainty through an uncertainty-aware visual data splitting strategy. Notably, CORAL introduces a false discovery rate (FDR) control mechanism that computes mirror statistics via symmetrically perturbed inputs and establishes data-driven thresholds, enabling training-free hallucination mitigation. The proposed approach is compatible with various mainstream architectures and significantly outperforms state-of-the-art methods across multiple benchmarks. By effectively suppressing hallucinations while preserving genuine object detection capabilities, this work substantially enhances the reliability and robustness of model outputs.
📝 Abstract
Multiple object hallucination, where large vision-language models (LVLMs) generate objects not supported by the visual input, is a persistent challenge caused by visual uncertainty during decoding. Existing methods reduce hallucinations using contrastive signals, but they rely on heuristics and lack principled control of false positives at the image level. To address this, we propose False Discovery Rate-COntRol of HALlucination (CORAL), a training-free framework that models visual uncertainty using an uncertainty-aware visual data splitting strategy and leverages mirror statistics to quantify visual contrast during decoding. By computing mirror statistics from paired, symmetrically perturbed visual inputs, CORAL estimates spurious object predictions and sets a data-driven threshold to control the expected fraction of false discoveries per image, suppressing hallucinations while retaining high power for truly grounded objects. The framework is flexible, supports multiple LVLMs, and mitigates hallucinations without retraining or supervision. Extensive experiments on multiple benchmarks with several evaluation metrics demonstrate that CORAL consistently outperforms state-of-the-art methods, providing more reliable and robust hallucination control. Code is available at: https://changliu1993-cl.github.io/CORAL/