๐ค AI Summary
This study addresses the limitations of costly human expert evaluations that hinder the deployment of large language models in high-stakes domains such as Arabic sociolinguistics and cultural understanding, where existing approaches struggle to balance linguistic fluency with deep cultural competence. Focusing on Egyptian and Iraqi Arabic dialects, the authors introduce the first weighted evaluation framework integrating both positive content criteria and negative error standards, grounded in 103 promptโrating pairs curated by native-speaking domain experts. To distinguish systematic bias from random noise, they propose two metrics: Mean Absolute Deviation (MAD) and Signed Mean Error. Experiments reveal that GPT-5.4 emerges as the strongest automatic rater (MADj = 10.21 percentage points), most models exhibit a systematic leniency bias, cultural tasks are inherently harder to score than linguistic ones, and implicit cultural reasoning constitutes the primary failure mode for automated evaluation.
๐ Abstract
The cost of human expert evaluation is a principal bottleneck to deploying language models in specialized, high-stakes domains. This is particularly acute for Arabic sociolinguistic knowledge: credible grading requires not only linguistic fluency but deep cultural familiarity that cannot be approximated by surface-level metrics. We address this with a cross-evaluation framework instantiated on two underrepresented Arabic dialect communities: Egyptian and Iraqi Arabic. We contribute 103 validated prompt-rubric pairs (70 Egyptian, 33 Iraqi; 53 Cultural, 50 Linguistic), authored and graded by native-speaker SMEs using penalty-weighted rubrics distinguishing positive content requirements from answer-specific negative error criteria. Three frontier LLMs serve as target models (graded by human SMEs across 302 unique prompt-response pairs), while five frontier LLMs serve as automated judges enforcing a provider-level self-evaluation guard. A dual-metric scheme combining Mean Absolute Deviation (MAD) with Signed Mean Error separates directional grading bias from symmetric noise. Across 1,307 judge evaluations: GPT-5.4 is the most reliable judge (MADj = 10.21 pp, Signed Error = -1.12%); four of five judges show systematic leniency (+2.01% to +6.56%); Cultural tasks are harder to grade than Linguistic tasks for all judges (MAD gap 1.83-4.78 pp); and models substantially outperform on Egyptian prompts compared to Iraqi prompts. However, given leniency differences between Iraqi and Egyptian SMEs, we cannot solely attribute this gap to model knowledge. We therefore emphasize findings that do not assume identical leniency across human graders. Across all samples, implicit cultural reasoning -- requiring models to simulate native-speaker judgment rather than rely on lexical verification -- emerges as the primary failure mode for automated grading across all judge models.