Beyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language Models

📅 2026-08-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
该研究针对阿拉伯语方言和文化能力评估不足的问题,提出了一种基于评分标准的沙特方言基准,并通过31个专家编写的提示来评测四个先进系统的性能。
📝 Abstract
Large language models are increasingly deployed in Arabic-speaking markets, yet standard benchmarks overwhelmingly reward Modern Standard Arabic (MSA) fluency while leaving dialectal and culturally grounded competence unmeasured. This gap is consequential: everyday Arabic is largely dialectal, and dialect encodes social meaning that MSA-centric evaluation cannot capture. We present a rubric-based benchmark for the Saudi dialect, comprising 31 expert-authored prompts spanning idiomatic, pragmatic, lexical, and culturally-embedded phenomena, each paired with an expert-established ground truth. Our methodology separates evaluation into a model-agnostic phase, in which atomic, MECE positive criteria are derived solely from the ground truth, and a model-specific phase, in which four state-of-the-art systems -- Claude Opus 5, Gemini 3.7, GPT-5.6, and Kimi K3 -- are scored against those criteria and penalised for errors they actively introduce. Across 124 model-prompt evaluations we catalogue 466 error instances under a nine-category taxonomy. The four systems cluster within a narrow macro-average band (42.7%-53.1%), with no model exceeding 55% and every model recording at least one negative-scoring prompt, confirming that Saudi dialectal competence remains broadly unsolved. Notably, Ambiguous Framing is the dominant failure mode (37.3% of errors) while outright Hallucination accounts for only 11.2%, indicating that models fail less by stating falsehoods than by distorting register and flattening pragmatic nuance. We further observe a consistency-versus-ceiling trade-off and model-distinctive error signatures. We release the full prompt set, ground truths, and scored rubrics to support reproducible dialectal evaluation.
Problem

Research questions and friction points this paper is trying to address.

Arabic Dialect
Cultural Competence
Large Language Models
Evaluation Benchmark
Modern Standard Arabic
Innovation

Methods, ideas, or system contributions that make the work stand out.

rubric-based benchmark
Saudi dialect
cultural competence
MECE criteria
dialectal evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
G
Ghassan Al-Sumaidaee
Perle
Sajjad Abdoli
Sajjad Abdoli
perle.ai
Deep LearningMusic Information RetrievalAdversarial Machine Learning
A
Ahmed Rashad
Perle
M
Maxim Legg
Perle