🤖 AI Summary
This study addresses the longstanding limitations in Arabic natural language processing (NLP), which stem not from linguistic complexity per se but from insufficient engagement with sociolinguistic realities and a lack of interdisciplinary integration. Through a systematic review of two decades of research, the work identifies critical gaps in infrastructure and paradigms, attributing them to inadequate social embedding and cross-disciplinary collaboration. By synthesizing efforts in language resource development, shared task organization, social media analysis, and computational social science, the project elucidates the challenges of transfer between Modern Standard Arabic and its dialects. It further underscores the sociocultural attributes of datasets and the pivotal role of task-oriented communities. The study advocates for a paradigm shift in low-resource NLP—one that transcends purely technical solutions to holistically address social, institutional, and cognitive dimensions—offering a new framework for global low-resource language research.
📝 Abstract
This paper reflects on twenty years of building NLP resources and research infrastructure for Arabic, a language spoken by hundreds of millions yet historically underserved relative to languages such as English or Chinese. The first decade focused on foundational linguistic infrastructure; the second shifted toward computational social science, social media analysis, and socially oriented applications. Rather than cataloguing outputs, the paper examines what the experience of building them revealed. Three counterintuitive lessons emerge: building datasets is as much a social process as a technical one; communities formed around shared tasks often matter more than the tasks themselves; and moving from language resources to computational social science exposes challenges that traditional NLP training does not address. We discuss three failures: a depression detection corpus that never reached clinical practice, a period of spreading across too many shared tasks without sufficient depth, and a long-standing assumption that Modern Standard Arabic infrastructure would transfer cleanly to dialectal tasks. These experiences suggest that the hardest problems in developing NLP for underserved communities are not linguistic but social, institutional, and epistemic, and require competencies the field rarely teaches.