Mitigating Accent-Language Confusion in Self-Supervised Speech Representations for Language Identification

πŸ“… 2026-10-07
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the tendency of self-supervised speech representations to conflate accent with language in language identification (LID), which frequently causes non-native speech to be misclassified as the speaker’s first language (L1). To mitigate this, we propose a training-free correction framework based on geometric projection. Without requiring second-language data or model fine-tuning, our method directly estimates and removes the L1 bias direction within the representation space of a frozen MMS-LID model, thereby eliminating accent interference at its source. Experimental results demonstrate that the proposed approach significantly improves target-language identification accuracy for non-native speech while maintaining stable predictive performance on native speech. This work establishes an efficient new paradigm for cross-accent robust LID.
πŸ“ Abstract
Spoken language identification (LID) aims to recognize the target language regardless of accent. In practice, however, LID models fine-tuned from self-supervised speech representations frequently confuse accents with languages, misclassifying non-native (L2) speech as the speaker's first language (L1). We show that non-native speech representations lie between native target-language and native L1 poles, causing systematic misclassification. To address this, we introduce a geometric projection that estimates an L1-bias direction solely from native speech and removes it before the frozen LID head. Across five MMS-LID models and non-native corpora, this projection substantially improves target language identification for L2-accented speech while preserving predictions for native speech. These results show that accent-induced L1 bias can be corrected directly within the representation space without L2 training data or model adaptation.
Problem

Research questions and friction points this paper is trying to address.

Spoken Language Identification
Self-Supervised Speech Representations
Accent-Language Confusion
Non-native Speech
L1 Bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Supervised Speech Representations
Language Identification
Geometric Projection
Accent-Language Confusion
L1 Bias Removal
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Minu Kim
Minu Kim
KAIST
speech recognitionspeaker verificationphonologylinguistics
Jihwan Lee
Jihwan Lee
PhD Student, Signal Analysis and Interpretation Lab (SAIL) at University of Southern California
brain-computer interfacesspeech synthesisbiosignal-to-speecharticulatory phonetics
D
David R. Mortensen
Language Technologies Institute, Carnegie Mellon University, USA
S
Shrikanth Narayanan
Signal Analysis and Interpretation Lab (SAIL), University of Southern California, USA