🤖 AI Summary
This study addresses the challenge of cross-lingual transfer in Arabic dialect classification by proposing BARRAC, a framework that adapts mainstream English sentiment analysis methods to this task. Methodologically, it replaces consumer review attribute pools with Arabic linguistic markers and employs a two-stage training strategy instead of noisy self-training, enabling effective transfer of aspect-based sentiment analysis and representation learning to dialect identification and sarcasm detection. Experimental results demonstrate that BARRAC achieves an average Macro-F1 score of 63.93% across five datasets, outperforming existing state-of-the-art methods by 3% and surpassing GPT-4o on most tasks.
📝 Abstract
With the rapid growth of Arabic NLP, several models, datasets and benchmarks have been reported. This paper asks whether approaches developed for majority languages like English can be adapted to Arabic tasks. We adapt an English aspect-based sentiment analysis framework to Arabic classification tasks and present the adaptation as BARRAC: Brainstorming Alignment and Replaced Representation learning for ArabiC tasks. BARRAC replaces consumer-review attribute pools with Arabic linguistic devices and markers for dialectal sentiment, sarcasm, and dialect identification, and replaces noisy self-training with two-stage training. Evaluated on five Arabic dialect datasets, BARRAC achieves a mean macro-F1 of 63.93\%, outperforming the best few-label SOTA by 3\%, and outperforming GPT-4o on four out of five tasks. Error analysis provides insights into remaining challenges. These results demonstrate that adapting task-specific approaches is a promising direction for Arabic NLP alongside adapting models, datasets and benchmarks.