Representation-based Broad Hallucination Detectors Fail to Generalize Out of Distribution

📅 2025-09-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
State-of-the-art hallucination detection methods exhibit severe out-of-distribution (OOD) generalization failure, heavily relying on spurious correlations present in training data and performing near-randomly on OOD benchmarks such as RAGTruth. Method: Through controlled ablation studies and supervised linear probing experiments, we systematically analyze the role of spurious correlations and propose “relevance robustness” as a new evaluation principle for hallucination detection. We further conduct cross-dataset generalization analysis to assess transferability and hyperparameter sensitivity. Contribution/Results: We demonstrate that removing spurious correlations yields no significant performance gain; remarkably, simple linear probes match the performance of complex models under relevance-robust evaluation. Our systematic evaluation reveals that mainstream methods require extensive dataset-specific tuning and suffer from poor cross-dataset transferability. This work provides both theoretical insights and empirical benchmarks for trustworthy evaluation and robust modeling of hallucination detection.

Technology Category

Machine Learning: Evaluation and AnalysisNatural Language Processing: Interpretability, Analysis, and Evaluation of NLP ModelsKnowledge Representation and Reasoning: Computational Complexity of Reasoning

Application Category

Web Mining and Content Analysis: Robustness and generalizability of Web mining methodsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
We critically assess the efficacy of the current SOTA in hallucination detection and find that its performance on the RAGTruth dataset is largely driven by a spurious correlation with data. Controlling for this effect, state-of-the-art performs no better than supervised linear probes, while requiring extensive hyperparameter tuning across datasets. Out-of-distribution generalization is currently out of reach, with all of the analyzed methods performing close to random. We propose a set of guidelines for hallucination detection and its evaluation.
Problem

Research questions and friction points this paper is trying to address.

Current hallucination detectors fail to generalize out-of-distribution
Performance is driven by spurious dataset correlations not true detection
All analyzed methods perform near random on unseen data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evaluated hallucination detectors' out-of-distribution generalization
Found performance driven by spurious data correlations
Proposed new evaluation guidelines for detection
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zuzanna Dubanowska
Samsung AI Research Center, Warsaw, Poland
M
Maciej Żelaszczyk
Samsung AI Research Center, Warsaw, Poland
Michał Brzozowski
Michał Brzozowski
Faculty of Economic Sciences, University of Warsaw
Paolo Mandica
Paolo Mandica
AI Research Scientist, Samsung AI Center, Warsaw
AI Safety and AlignmentGenerative AIUncertainty EstimationHyperbolic Neural Networks
M
Michał Karpowicz
Samsung AI Research Center, Warsaw, Poland