The Detectability Gap: Hidden Heterogeneity in Hallucination Detection Across Language Models

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses how aggregated performance metrics obscure systematic disparities in detecting large language model (LLM) hallucinations via sampling consistency. Focusing on this hidden heterogeneity, we propose a taxonomy classifying hallucinations as “ghosts” and “flickers,” employing lexical and semantic response dispersion analysis alongside bootstrap statistical testing. Our results reveal a substantial detectability gap of 0.35 to 0.46 AUC between high- and low-consistency hallucinations, demonstrating that conventional aggregated metrics mask persistent, model-dependent failure modes. By exposing these overlooked structural biases in hallucination detection, this work provides critical empirical foundations for developing more robust evaluation frameworks.
📝 Abstract
Sampling based consistency is widely used for hallucination detection, yet aggregate performance can conceal systematic differences in which errors are detectable. This work studies that heterogeneity across four language models and three factual question answering datasets. Partitioning hallucinations by answer agreement reveals high agreement (Ghost) and low agreement (Flickering) regimes with an apparent detectability gap of $0.35$ to $0.46$ AUC. Because the statistics used to define the regimes and measure this gap are strongly coupled ($|ρ|\approx0.94$ to $1.00$), the raw result is treated as a property of agreement based detection rather than independent evidence. After freezing regime assignments, lexical and semantic response dispersion preserve the asymmetry, with bootstrap $95\%$ intervals excluding zero in all $12$ model and dataset settings. A stricter test using individual diffusion trajectories and no cross seed information preserves the asymmetry across all three LLaDA datasets ($p<0.005$) and directionally across all three Dream datasets, with one reaching significance. The hard regime varies substantially in prevalence across models ($16\%$ to $77\%$), and matched prompts frequently change regimes between models. These findings show that aggregate detection metrics conceal persistent, model dependent heterogeneity in language model failures and motivate regime conditioned evaluation.
Problem

Research questions and friction points this paper is trying to address.

Hallucination Detection
Language Models
Detectability Gap
Hidden Heterogeneity
Sampling-based Consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hallucination Detection
Sampling-based Consistency
Detectability Gap
Heterogeneity
Regime-conditioned Evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Pranav Darshan
Department of Computer Science and Engineering, R.V. College of Engineering, India
P
Pranav A
Department of Computer Science and Engineering, R.V. College of Engineering, India
S
Sravan Karthick T
Department of Computer Science and Engineering, R.V. College of Engineering, India
M
Minal Moharir
Department of Computer Science and Engineering, R.V. College of Engineering, India
Ivan P. Yamshchikov
Ivan P. Yamshchikov
Research Professor at CAIRO, THWS
natural language generationcomputational creativityempathetic aiethics of ai application