🤖 AI Summary
This study addresses the generalization challenges of hallucination detection and fact verification in Arabic large language models under distribution shifts. To this end, we introduce HalluScore, the first large-scale evaluation benchmark tailored for Arabic, accompanied by a shared task. Methodologically, we establish rigorous cross-model and cross-question generalization settings, conducting systematic evaluations using the HalluTruthQA dataset, the AUC-ROC metric, and multi-candidate answer ranking techniques. The shared task attracted 13 participating teams. The best-performing system achieved an AUC of 0.77 for binary hallucination detection and a fact verification score exceeding 0.85. These results effectively reveal the performance bottlenecks and limitations of current approaches in low-resource language scenarios.
📝 Abstract
We present HalluScoring 2026, a shared task for evaluating hallucination detection and factual verification in Arabic question answering under challenging generalization settings. The shared task is organized into two main tasks, each comprising two subtasks, for a total of four subtasks. Task 1 evaluates binary hallucination detection, considering generalization to unseen questions (Subtask 1.1) and responses generated by unseen LLMs (Subtask 1.2). Task 2 extends the evaluation beyond detection by requiring the systems to additionally identify the correct factual answer from six related candidates, covering Islamic knowledge (Subtask 2.1) and general knowledge (Subtask 2.2). The shared task is based on two Arabic datasets: HalluScore and HalluTruthQA. A total of 13 teams participated in the shared task, 10 of which submitted system description papers. The results of Task 1 demonstrate that hallucination detection remains challenging under distribution shift, with the winning team achieving AUC-ROC test scores of 0.772 and 0.767 for Subtasks 1.1 and 1.2, respectively. For Task 2, the winning team achieved scores of 0.882 and 0.857 in the Islamic and general-knowledge subtasks, respectively, under assisted evaluation.