Different Perturbations, Different Mechanisms: Understanding Continued Pre-training for Zero-Shot Dialect Robustness

๐Ÿ“… 2026-08-05
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Multilingual large language models often underperform on dialectal variants, and the mechanisms underlying the effectiveness of perturbation-based continued pretraining (CPT) remain poorly understood. This study systematically evaluates six perturbation strategies across nine dialectal tasks in German, Italian, and Arabic, combining representational analysis with predictive behavior assessment. The findings reveal that perturbation methods yielding comparable performance improvements enhance zero-shot dialect robustness through distinct mechanismsโ€”such as language model adaptation, representation alignment, and prediction repair. Notably, character-noise CPT substantially improves generalization to dialects without compromising performance on standard language varieties. These results provide empirical guidance and theoretical insights for selecting effective CPT strategies in multilingual settings.
๐Ÿ“ Abstract
Dialectal variation remains a major challenge for multilingual language models. Perturbation-based continued pre-training (CPT) has emerged as a promising approach to improving robustness, yet existing work largely evaluates individual perturbation strategies in isolation and provides limited insight into why they work. We present a systematic study of perturbation-based CPT for multilingual dialect robustness in LLMs, comparing six training conditions across nine German, Italian, and Arabic dialect tasks. Perturbation-based CPT, especially character-noised CPT, consistently improves zero-shot dialect robustness while largely preserving standard variety performance. More importantly, we show that methods with similar downstream performance induce distinct mechanisms of robustness, exhibiting different patterns of language model adaptation, representational alignment, and prediction repair. Our results provide a more complete understanding of how synthetic surface variation improves robustness and offer practical guidance for selecting CPT strategies in multilingual and dialectal settings.
Problem

Research questions and friction points this paper is trying to address.

dialect robustness
multilingual language models
zero-shot
perturbation-based continued pre-training
dialectal variation
Innovation

Methods, ideas, or system contributions that make the work stand out.

perturbation-based continued pre-training
dialect robustness
zero-shot transfer
representational alignment
character-noised CPT
๐Ÿ”Ž Similar Papers
2024-01-31Workshop on Representation Learning for NLPCitations: 3