AraDynFact: Dynamic Evaluation of Factual Knowledge in Arabic

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing factual knowledge evaluations for Arabic large language models (LLMs), which predominantly rely on costly manual curation and lack cultural sensitivity. To overcome these challenges, this work proposes a fully automated, dynamic evaluation paradigm that eliminates human intervention. The framework leverages data mining from Arabic Wikipedia, integrating dynamic information extraction with automatic question-answering generation to comprehensively audit model performance across historical, societal, and regional dimensions. By successfully evaluating multiple state-of-the-art models, the results demonstrate that this approach achieves high consistency with existing manually constructed benchmarks. Ultimately, this research establishes a cost-effective and highly scalable pathway for assessing factual knowledge in low-resource languages.
📝 Abstract
As Large Language Models (LLMs) continue to scale both in size and capabilities, their proficiency in the Arabic Language has seen significant advancement. However, a critical gap remains: the extent of their factual knowledge and cultural sensitivity to the diverse Arabic-speaking world remains largely underexplored. Current evaluation metrics often focus on translation or generic reasoning, failing to capture the rich historical, social, and regional nuances inherent to Arabic culture. In addition, most benchmarks rely on heavy work, with human intervention in some steps, making the evaluation of knowledge coverage expensive and slow. To address this deficiency, we introduce AraDynFact, a novel dynamic evaluation framework designed to rigorously assess the factual Arabic knowledge embedded in LLMs. Unlike static benchmarks, AraDynFact employs a dynamic approach to extract factual information and generate rich and answerable questions in a fast and automatic way. We apply AraDynFact to Arabic Wikipedia and audit the performance of several state-of-the-art models, ranging from Arabic-centric specialized LLMs to high-resource general purpose LLMs. In addition we found a high degree of correlation with existing, hand-crafted Arabic-centric benchmarks, confirming the potential of our dynamic approach.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Arabic factual knowledge
Cultural sensitivity
Dynamic evaluation
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic Evaluation Framework
Factual Knowledge Assessment
Arabic LLMs
Automated Question Generation
Cultural Sensitivity
🔎 Similar Papers
No similar papers found.