Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment

📅 2026-05-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing benchmarks in evaluating large language models’ (LLMs’) ability to detect contextual hallucinations, which often rely on single human annotations and thus struggle with semantic ambiguity, potentially leading to underestimation of model performance. To mitigate this, the authors propose a hybrid evaluation paradigm that prioritizes LLMs followed by human arbitration: initial hallucination predictions—both at the rationale and span levels—are generated by GPT-5 Mini and Gemini 2.5 Flash, after which conflicting cases undergo dual, cross-cultural human review. This approach substantially improves annotation consistency, increasing three-way agreement by 6.38%–7.62%, corrects biases in the original assessment, and boosts detection accuracy for GPT and Gemini by 2.34%–4.25% and 3.80%–8.51%, respectively, with inter-annotator agreement among human arbiters reaching 83%–87%.
📝 Abstract
Hallucination remains a persistent challenge in Large Language Models (LLMs), particularly in context-grounded settings such as RAG and agentic AI systems. This study focuses on contextual hallucination detection in summarization tasks. We analyze the QAGS-C and SummEval datasets by comparing original benchmark annotations with reason and span-based predictions from Gemini 2.5 Flash and GPT-5 Mini. To address systematic divergences between human labels and LLM judgments, we re-evaluated all conflicted samples through a human adjudication process involving 2 cross-cultural adjudicators. Following this re-evaluation, triple agreement (between human, GPT, and Gemini) increased by 6.38% for QAGS-C and 7.62% for SummEval. Similarly, model accuracy improved, with GPT increasing by 4.25% on QAGS-C and 2.34% on SummEval, while Gemini showed gains of 8.51% and 3.80%, respectively. Notably, adjudicators frequently sided with the models' judgments over original human annotations when LLMs provided explicit reasoning. Overall human adjudicator agreement ranged between 83% and 87%. These findings suggest that for ambiguity-prone tasks, single-pass annotations may be insufficient, and model-assisted re-evaluation yields more reliable benchmarks.
Problem

Research questions and friction points this paper is trying to address.

hallucination detection
benchmark evaluation
large language models
human annotation
summarization
Innovation

Methods, ideas, or system contributions that make the work stand out.

hallucination detection
LLM-first evaluation
human adjudication
benchmark reliability
contextual summarization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
I
Ismail Furkan Atasoy
Department of Computer Engineering, Hacettepe University, Ankara, 06800, Türkiye
B
Begum Mutlu
Department of Computer Engineering, Ankara University, Ankara, 06830, Türkiye
Ebru Akcapinar Sezer
Ebru Akcapinar Sezer
Hacettepe University
AIData QualityModelingSearchingSemantics
A
Abdelrahman Wahdan
Zephlen AI and Information Technologies Inc., Kocaeli, 41400, Türkiye