The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the annotation bias in existing hallucination detection benchmarks caused by conflating reference faithfulness with factual correctness. To investigate this, we construct a dataset of 900 human-annotated instances and conduct comparative experiments using controlled prompt variants alongside lexical similarity, natural language inference (NLI), and LLM-as-a-judge techniques to systematically evaluate automatic annotation strategies under the factual correctness objective. Our findings reveal that the selection of label sources is central to benchmark design, demonstrating that optimized prompts significantly enhance human–machine agreement while reducing false positive rates. Furthermore, this work underscores the necessity of explicitly verifying and ensuring strict alignment between annotation criteria and evaluation objectives, thereby offering critical guidance for the standardized design of hallucination detection benchmarks.
📝 Abstract
In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed. These methods are often benchmarked with open-domain question answering (QA) datasets containing questions and corresponding short reference answers. First, an LLM is used to generate answers to questions within the QA dataset. Then, some automated labeling strategy is used to label these answers as hallucinated or not by comparing them with the reference answers in the dataset. This evaluation setting creates a methodological ambiguity between two criteria: reference faithfulness (whether the answer is fully supported by the reference) and factual correctness (whether the answer is free from contradictions and factually false specific claims). In practice, automated labelers may apply the former criterion even when the intended target is the latter. We study this potential criterion mismatch using 900 human-labeled question-answer pairs spanning three commonly used QA datasets and three generator models, with labels targeting answer-level factual correctness. We evaluate lexical similarity metrics, a reference-entailment NLI baseline, and seven LLM judges under controlled prompt variants as automated labelers. Our experiments reveal substantial disagreement both among automated labeling strategies and between these labels and human annotations. Many strategies also exhibit strong directional error biases, and for most judge-generator pairs, replacing a faithfulness-oriented prompt with a factual-correctness prompt improves agreement with human annotations and reduces false-positive dominance, indicating that automated hallucination labels depend strongly on how the target criterion is specified. Label-source choice should therefore be considered a fundamental part of benchmark design and made explicit, validated, and matched with the benchmark goal.
Problem

Research questions and friction points this paper is trying to address.

hallucination detection
labeling problem
benchmark evaluation
factual correctness
reference faithfulness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hallucination Detection
Labeling Problem
Factual Correctness
Reference Faithfulness
LLM Judges
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jorma Valjakka
Department of Computer Science, University of Helsinki
J
Juhani Kivimäki
Department of Computer Science, University of Helsinki
J
Juha Mylläri
Department of Computer Science, University of Helsinki
Jukka K. Nurminen
Jukka K. Nurminen
Professor of Computer Science, University of Helsinki
practice of AIquantum softwaremobile systemsenergy-efficiency#UnivHelsinkiCS