Phonological Interference in Multilingual Speech Models

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the pronunciation errors in multilingual speech models arising from enforced monolingual phonology during code-switching or low-resource scenarios. Through internal activation probing, we reveal that such interference originates within low-dimensional subspaces. To mitigate this, we introduce Windowed Language Estimation (WLE), a novel technique that dynamically corrects local language hypotheses during inference to preserve accurate phonemes without requiring retraining. Our approach eliminates 34%–69% of phonological interference while maintaining lossless monolingual performance. Furthermore, it significantly enhances transcription and generation quality for both mixed and unseen languages, offering an efficient, training-free robustness enhancement solution for multilingual speech processing.
📝 Abstract
Phoneme-level models transcribe or generate speech as a sequence of phonemes, the smallest sound units that distinguish words. These models enable fine-grained pronunciation control and understanding, yet often fail on input that does not match any single training language, such as speech alternating between two languages, known as code-switching, or low-resource languages absent from training. We identify a systematic failure mode behind this, phonological interference: models assume the input is in a single language and impose its phonology, overriding local phoneme-level decisions that conflict with the assumed language. We measure interference by how often a model retains phonemes that one language has but the other lacks. On code-switched input, two phone recognizers (speech-to-phoneme models) and a phoneme-conditioned text-to-speech model lose 32% to 79% of these phonemes, but lose far fewer of the phonemes both languages share. On unseen languages, we find that phone recognizers impose the phonology of the training language they assign to the speech, and the more confident the assignment, the more they lose phonemes the unseen language has but the assigned language lacks. We probe the models' language estimate from their internal activations, and trace interference to a low dimensional subspace. On monolingual speech, steering this subspace toward another language makes the model lose the phonemes that only the original language uses and produce phonemes that only the target language has. We introduce windowed language estimation (WLE), an inference time repair that replaces the model's language estimate in this subspace with one computed from a short window around each position. On code-switched input, WLE removes 34% to 69% of the interference in all three models, and in the recognizers it leaves monolingual performance essentially unchanged.
Problem

Research questions and friction points this paper is trying to address.

Phonological Interference
Multilingual Speech Models
Code-switching
Phoneme-level models
Low-resource languages
Innovation

Methods, ideas, or system contributions that make the work stand out.

Phonological Interference
Windowed Language Estimation
Code-switching
Low-dimensional Subspace
Multilingual Speech Models
🔎 Similar Papers
No similar papers found.