Code-Switching Spoken Language Identification as Multi-Label Set Prediction

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited recognition robustness caused by code-switched speech leaking into monolingual filters. We formulate code-switching language identification (CS-LID) as a multi-label set prediction task and propose a model that directly generates language combinations, thereby overcoming the limitations of conventional approaches that rely on predefined language counts and threshold-based separation. Experimental results demonstrate that the proposed method accurately predicts the number of languages in unseen language pairs without requiring prior specification. Although its exact set accuracy remains slightly below that of an oracle top-k baseline, our approach effectively reveals the distributional gap between synthetic and real-world data, highlighting a core bottleneck in CS-LID research.
📝 Abstract
Code-switched (CS) speech leaks through the monolingual language identification (LID) filters used to curate massive speech corpora, calling for CS-aware LID (CS-LID). We formulate utterance-level CS-LID as multi-label language-set prediction and propose a set generator that directly outputs the languages in an utterance, comparing it against atomic-pair and score-based classification baselines. Oracle Top-k is the strongest baseline, but thresholding fails because no single threshold separates CS from monolingual speech. Our set generator predicts the correct language count on unseen pairs without assuming the number of languages, but underperforms oracle Top-k in exact set accuracy. Our analysis identifies the key obstacles to robust CS-LID: oracle cardinality, threshold instability, language bias in CS training data, and the synthetic-to-real gap.
Problem

Research questions and friction points this paper is trying to address.

Code-Switching
Spoken Language Identification
Multi-Label Set Prediction
Threshold Instability
Synthetic-to-Real Gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

Code-Switching
Spoken Language Identification
Multi-Label Set Prediction
Set Generator
Synthetic-to-Real Gap
🔎 Similar Papers
No similar papers found.