🤖 AI Summary
Arabizi—the Latin-script transliteration of Arabic—has long been overlooked in natural language processing, with limited understanding of its cross-dialectal usage patterns and systematic variation. This study presents the first large-scale, human-centered cross-dialectal investigation, encompassing Algeria, Egypt, Lebanon, Morocco, and Tunisia, to examine user perceptions and orthographic conventions in Arabizi writing. Through character-level alignment, we analyze lexical transliteration variation and construct a manually curated sentence-level parallel corpus of Arabic and Latinized text. Our contributions include the first character-aligned lexicon spanning these five dialects and a high-quality parallel sentence corpus, demonstrating that Arabizi exhibits structured, systematic properties rather than ad hoc improvisation. These resources provide a critical empirical foundation and essential data for future NLP research on Arabizi.
📝 Abstract
Arabizi refers to Arabic written in Latin script. Although previous studies have shown that the prevalence and usage of Arabizi vary by factors such as region and age group, most NLP research on Arabic texts treats it as a temporary phenomenon resulting from limited technological support for the Arabic script. In this work, we engage with Arabic speakers to collect insights on their perceptions and usage of Arabizi. We further examine writing norms among speakers of different dialects, focusing on Algerian, Egyptian, Lebanese, Moroccan, and Tunisian Arabic. To this end, we release two resources. First, a character-level alignment of Arabic words to study inter- and intra-dialectal variation across these five dialects, based on words transliterated by survey participants, finding systematic intra-dialectal regularity and inter-dialectal variation. Second, to study Arabic speakers' ability to identify this stylistic variation at the sentence-level, we build a manually curated parallel corpus of sentences written in Arabic script alongside multiple Arabizi transliterations, collected from speakers of the same five dialects. Our study presents the largest human-centered, cross-dialectal study of Arabizi's perceptions and practices to date.