Does Modern Standard Arabic (MSA) Dominate Arabic Dialects in LLMs? A Representation-Level Analysis

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the internal representational mechanisms underlying the tendency of large language models to default to Modern Standard Arabic when generating Arabic dialects. Spanning 26 Arabic varieties, the research employs a language dominance framework, normalized mutual information, and multi-level, cross-model comparative analyses to examine whether internal representations are dominated by the standard variety and to disentangle generative biases from geometric structures. The findings challenge prevailing assumptions: output preferences do not equate to representational dominance. Specifically, Modern Standard Arabic does not dominate model internals; instead, dialectal representation spaces exhibit dense overlap, with separability varying significantly across architectures. This work elucidates fundamental properties of multilingual representations in large language models, offering novel perspectives for processing low-resource dialects.
📝 Abstract
Large language models (LLMs) often default to Modern Standard Arabic (MSA) when generating Arabic, even when prompted with dialectal Arabic. A natural explanation is that their internal representations are dominated by MSA. We test this hypothesis by adapting the language-dominance framework of Shani and Basirat (2025) (https://doi.org/10.18653/v1/2025.blackboxnlp-1.7) to 26 Arabic varieties. Across layers and model families, we find no evidence that MSA acts as a dominant internal representation for Arabic dialects. Instead, dialect representations form a dense and highly overlapping space: normalized mutual information drops sharply compared to patterns reported for more distinct languages. Moreover, the strongest separability effects are not confined to intermediate layers, but can shift toward later layers depending on the architecture. These findings challenge a common interpretation of MSA-biased generation: output preference does not necessarily reveal internal representational dominance. Analyses of multilingual and dialectal LLMs should therefore distinguish generation bias from the geometry of internal representations.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Modern Standard Arabic
Arabic Dialects
Internal Representations
Generation Bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

Representation-level analysis
Arabic dialects
Modern Standard Arabic
Normalized mutual information
Generation bias