Learning the Arabic Dialect Continuum as a Continuous Space: A Regression Approach to Speaker Origin Prediction

📅 2026-07-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of traditional Arabic dialect geolocation approaches, which treat dialects as discrete categories and fail to capture their inherent continuous geographic variation. To overcome this, the authors propose the first end-to-end continuous geolocation regression framework that models dialects as a geographic continuum. The model integrates speech representations from XLS-R-300M and Whisper-large-v3 with phonological features, employing a Transformer architecture and learnable attention pooling to directly predict speaker coordinates. A spherical geodesic loss is introduced to optimize great-circle distance. Evaluated under a strict no-data-leakage 5-fold GroupKFold protocol, the approach achieves a median localization error of 481.2 km, with country- and city-level classification accuracies of 64.5% and 45.2%, respectively. In zero-shot city-masking tests, the error increases to 1173.3 km.
📝 Abstract
We present a regression-based approach to Arabic dialect geolocation that models dialectal variation as a continuous geographic space rather than discrete categories. Speaker origin is predicted as continuous latitude-longitude coordinates using a hierarchical neural architecture that fuses frame-level XLS-R-300M and Whisper-large-v3 encoder representations with phonotactic descriptors through a Transformer encoder and a learnable attention-pooled query. A spherical geodesic loss directly optimizes great-circle distance on Earth's surface, avoiding distortions inherent to planar coordinate regression. Under a leakage-free 5-fold GroupKFold protocol grouped by source recording, our model attains a pooled median localization error of 481.2 km. Auxiliary country and city heads reach 64.5% and 45.2% accuracy, respectively. A permutation Mantel test on the learned latent space provides quantitative support for the Arabic dialect continuum hypothesis. To probe true generalization, we further introduce a city-masking protocol in which two cities per fold are removed from training but retained in validation. Under this zero-shot regime, the mean error rises to 1173.3 km, a 1.32x degradation relative to seen cities. Our findings establish continuous geographic modeling as a principled framework for Arabic dialect geolocation and quantify both its strengths and the substantial headroom that remains.
Problem

Research questions and friction points this paper is trying to address.

Arabic dialect continuum
geolocation
continuous space
speaker origin prediction
dialectal variation
Innovation

Methods, ideas, or system contributions that make the work stand out.

continuous dialect modeling
spherical geodesic loss
hierarchical neural fusion
zero-shot city masking
Arabic dialect continuum
🔎 Similar Papers
No similar papers found.