HalluTruthQA-4K: A Fine-Grained Corpus and Annotation Process for Arabic Hallucination Detection and Truth Verification

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of fine-grained annotations in existing Arabic hallucination detection datasets, which typically provide only coarse binary labels without specifying error locations, causes, or correct facts. To bridge this gap, the authors introduce the first fine-grained hallucination detection dataset for Arabic, spanning four domains—Islamic knowledge, history, science, and geography—and comprising 4,000 expert-annotated question-answer pairs. Each instance includes a model-generated response, a reference answer, and high-quality distractors, along with novel character-level hallucinated span annotations, human-written explanations, and a hierarchical hallucination taxonomy. Rigorous data quality is ensured through controlled generation, multi-round expert annotation, and arbitration. The dataset contains 1,643 hallucinated and 2,357 non-hallucinated responses, with 1,843 annotated erroneous spans, and has been adopted as the official benchmark for the HalluScoring 2026 shared task.
📝 Abstract
Large language models can generate fluent Arabic answers while introducing factual errors that are difficult to identify and verify. Existing Arabic hallucination resources often assign a binary label to an entire response, indicating whether it is hallucinated or non-hallucinated, but provide limited information about the exact erroneous content, the reason for the error, or the correct factual answer. We present HalluTruthQA-4K, an expanded version of the HalluTruthQA resource containing 4,000 expert-curated Arabic question-answering instances across four knowledge-intensive domains: Islamic knowledge, history, science, and geography. Serving as the official dataset for Track 2 of the HalluScoring 2026 shared task, HalluTruthQA-4K extends our original corpus to 4,000 instances. Each instance pairs an Arabic question with a model-generated response, a verified reference answer, and five plausible distractors. Hallucinated responses are additionally annotated with character-level erroneous spans, human-written explanations, and hierarchical hallucination types. The corpus contains 1,643 hallucinated and 2,357 non-hallucinated responses, with 1,843 annotated erroneous spans. We describe the resource construction and annotation methodology, including question selection, controlled answer generation, candidate construction, expert annotation, independent verification, adjudication, and quality control. We also document the annotation guidelines, taxonomy, data format, inter-annotator agreement, and corpus statistics. HalluTruthQA-4K provides a reusable resource for hallucination detection, span-level error localization, explanation generation, factual verification, and the broader evaluation of factual reliability in Arabic language models.
Problem

Research questions and friction points this paper is trying to address.

hallucination detection
fact verification
Arabic language models
fine-grained annotation
truthfulness evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

fine-grained annotation
hallucination detection
Arabic language models
fact verification
error span localization
🔎 Similar Papers
No similar papers found.