Halluscoring 2026: The first shared task on llms hallucination detection and answer verification

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the generalization challenges of hallucination detection and fact verification in Arabic large language models under distribution shifts. To this end, we introduce HalluScore, the first large-scale evaluation benchmark tailored for Arabic, accompanied by a shared task. Methodologically, we establish rigorous cross-model and cross-question generalization settings, conducting systematic evaluations using the HalluTruthQA dataset, the AUC-ROC metric, and multi-candidate answer ranking techniques. The shared task attracted 13 participating teams. The best-performing system achieved an AUC of 0.77 for binary hallucination detection and a fact verification score exceeding 0.85. These results effectively reveal the performance bottlenecks and limitations of current approaches in low-resource language scenarios.
📝 Abstract
We present HalluScoring 2026, a shared task for evaluating hallucination detection and factual verification in Arabic question answering under challenging generalization settings. The shared task is organized into two main tasks, each comprising two subtasks, for a total of four subtasks. Task 1 evaluates binary hallucination detection, considering generalization to unseen questions (Subtask 1.1) and responses generated by unseen LLMs (Subtask 1.2). Task 2 extends the evaluation beyond detection by requiring the systems to additionally identify the correct factual answer from six related candidates, covering Islamic knowledge (Subtask 2.1) and general knowledge (Subtask 2.2). The shared task is based on two Arabic datasets: HalluScore and HalluTruthQA. A total of 13 teams participated in the shared task, 10 of which submitted system description papers. The results of Task 1 demonstrate that hallucination detection remains challenging under distribution shift, with the winning team achieving AUC-ROC test scores of 0.772 and 0.767 for Subtasks 1.1 and 1.2, respectively. For Task 2, the winning team achieved scores of 0.882 and 0.857 in the Islamic and general-knowledge subtasks, respectively, under assisted evaluation.
Problem

Research questions and friction points this paper is trying to address.

Hallucination detection
Answer verification
Arabic question answering
Large language models
Generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hallucination Detection
Arabic QA
Generalization
Factual Verification
Distribution Shift
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Aisha Alansari
Aisha Alansari
Graduate Assistant, Information and Computer Science Department, KFUPM
Machine LearningNatural Language ProcessingDeep LearningLLMs
A
Abdessalam Bouchekif
Hamad Bin Khalifa University, Qatar
A
Ahmed Hasanaath
King Fahd University of Petroleum and Minerals, Saudi Arabia
Salah Eddine Bekhouche
Salah Eddine Bekhouche
University of the Basque Country
Face Analysis
M
Malak Alkhorasani
Imam Abdulrahman bin Faisal University, Saudi Arabia
M
Mohammed-En-Nadhir Zighem
Hamad Bin Khalifa University, Qatar
S
Saad Ezzini
King Fahd University of Petroleum and Minerals, Saudi Arabia
H
Hichem Telli
University of Biskra, Algeria
H
Hend Al-Khalifa
King Saud University, Saudi Arabia
Muhammad Abdul-Mageed
Muhammad Abdul-Mageed
The University of British Columbia
Natural Language ProcessingDeep Learning
H
Hadid Abdenour
Universiti Malaysia Kelantan, Malaysia
Hamzah Luqman
Hamzah Luqman
Associate Professor, King Fahd university for Petroleum and Minerals (KFUPM)
Computer VisionArabic Natural Language Processing