🤖 AI Summary
This work addresses the challenge in multilingual speaker verification where language-specific acoustic variations often entangle speaker identity with linguistic characteristics, thereby degrading cross-lingual generalization. To mitigate this issue, the authors propose L-Proto, a novel approach that introduces a language-aware prototypical training mechanism. By enforcing language-consistent sampling at the utterance level—restricting each training episode to utterances from a single language—the method explicitly controls for linguistic variability, effectively disentangling speaker and language representations. L-Proto is compatible with various backbone architectures and consistently outperforms conventional fine-tuning and random sampling strategies on the TidyVoice Challenge benchmark, achieving significant and robust performance gains across different model designs.
📝 Abstract
Multilingual speaker verification remains challenging because language-dependent acoustic variability causes speaker identity to become entangled with linguistic characteristics, degrading generalization across languages. In multilingual training, embeddings often encode language cues with speaker identity, causing speakers to form language-specific clusters. We propose L-Proto, a language-aware episodic prototypical training strategy that constructs language-consistent episodes. By sampling speakers from a single language per episode, L-Proto reduces language-driven variation during training and encourages embeddings to focus more directly on speaker identity. Experiments on the TidyVoice Challenge benchmark demonstrate consistent performance improvements over conventional fine-tuning and random episodic sampling across multiple backbone architectures.