🤖 AI Summary
This study addresses hallucination detection in black-box large language model (LLM) outputs. We propose a training-free, zero-shot, multilingual approach that quantifies token-level response uncertainty via entropy variation across multiple stochastic sampling runs. Our method localizes hallucinations at fine granularity by analyzing response consistency and optimizing hyperparameters for entropy-based confidence estimation. Crucially, it requires no model fine-tuning or labeled data, ensuring low computational cost, strong cross-lingual generalizability, and broad compatibility with arbitrary black-box LLMs. Evaluated on the SemEval-2025 Task 3 (Mu-SHROOM) multilingual hallucination detection benchmark, our approach achieves state-of-the-art localization accuracy. Error analysis further uncovers recurrent hallucination patterns and exposes fundamental limitations of current LLMs—particularly in factual grounding, logical coherence, and cross-lingual knowledge transfer. The framework thus provides both a practical, deployable tool for hallucination auditing and novel insights into LLM reliability boundaries.
📝 Abstract
Identification of hallucination spans in black-box language model generated text is essential for applications in the real world. A recent attempt at this direction is SemEval-2025 Task 3, Mu-SHROOM-a Multilingual Shared Task on Hallucinations and Related Observable Over-generation Errors. In this work, we present our solution to this problem, which capitalizes on the variability of stochastically-sampled responses in order to identify hallucinated spans. Our hypothesis is that if a language model is certain of a fact, its sampled responses will be uniform, while hallucinated facts will yield different and conflicting results. We measure this divergence through entropy-based analysis, allowing for accurate identification of hallucinated segments. Our method is not dependent on additional training and hence is cost-effective and adaptable. In addition, we conduct extensive hyperparameter tuning and perform error analysis, giving us crucial insights into model behavior.