π€ AI Summary
This study addresses the tendency of self-supervised speech representations to conflate accent with language in language identification (LID), which frequently causes non-native speech to be misclassified as the speakerβs first language (L1). To mitigate this, we propose a training-free correction framework based on geometric projection. Without requiring second-language data or model fine-tuning, our method directly estimates and removes the L1 bias direction within the representation space of a frozen MMS-LID model, thereby eliminating accent interference at its source. Experimental results demonstrate that the proposed approach significantly improves target-language identification accuracy for non-native speech while maintaining stable predictive performance on native speech. This work establishes an efficient new paradigm for cross-accent robust LID.
π Abstract
Spoken language identification (LID) aims to recognize the target language regardless of accent. In practice, however, LID models fine-tuned from self-supervised speech representations frequently confuse accents with languages, misclassifying non-native (L2) speech as the speaker's first language (L1). We show that non-native speech representations lie between native target-language and native L1 poles, causing systematic misclassification. To address this, we introduce a geometric projection that estimates an L1-bias direction solely from native speech and removes it before the frozen LID head. Across five MMS-LID models and non-native corpora, this projection substantially improves target language identification for L2-accented speech while preserving predictions for native speech. These results show that accent-induced L1 bias can be corrected directly within the representation space without L2 training data or model adaptation.