🤖 AI Summary
This work addresses the absence of high-quality, multitask evaluation benchmarks for assessing large language models’ comprehension and reasoning capabilities in the domain of Islamic traditional scholarship (turath). It presents the first systematically constructed, expert-developed and vetted benchmark that spans the full scope of this intellectual tradition, encompassing 3,465 question-answer pairs drawn from 35 classical texts dating from the 12th century onward, across seven disciplinary areas. Organized along dual dimensions of academic difficulty and task type, the benchmark integrates multiple formats—including multiple-choice, passage comprehension, and open-ended questions—and employs both expert annotations and a zero-shot evaluation framework to enable fine-grained analysis of model performance within deep historical and cultural contexts. The study also provides zero-shot baselines for ten prominent large language models alongside reference scores from domain scholars, establishing a foundational resource for evaluating specialized model competence in religious and cultural domains.
📝 Abstract
Large language models (LLMs) are increasingly used for question answering, education, and research, including in religious and cultural domains where answers depend on specialised source traditions. Yet in Islamic Studies, key concepts, methods, and debates preserved in the authoritative scholarly tradition, known as turath, lack high-quality annotated resources. We introduce IslamicTurathBench (ISTB), a multi-task, multi-discipline dataset for evaluating LLMs on classical Islamic scholarship. Developed and reviewed by domain experts, ISTB contains 3,465 question-answer items drawn from 35 recognised source works spanning more than 12 centuries of scholarship across seven key fields of Islamic Studies. To enable comprehensive profiling of model capabilities, ISTB is structured along two axes: scholarly demand (Beginner, Intermediate, and Advanced) and task format (multiple-choice questions, passage-based comprehension, and open-ended knowledge questions). ISTB includes aggregated scores from a scholarly human reference panel and zero-shot baselines from ten systems. The dataset supports reproducible evaluation of language-model behaviour across source works, disciplines, scholarly-demand levels, and question formats in a historically layered scholarly domain.