🤖 AI Summary
This study addresses the challenge of forced alignment in Hindi–English code-switched speech, where phonetic free variation and ambiguous intra-sentential English word boundaries degrade alignment performance. To overcome this, the authors propose a novel approach that integrates guided lexicon expansion with acoustic model fine-tuning specifically tailored for code-switched data. By extending the Montreal Forced Aligner (MFA) dictionary and fine-tuning the acoustic model on authentic code-switched corpora, the method substantially improves alignment accuracy. Experimental results demonstrate that the proposed approach reduces the average alignment error to 4.15 milliseconds—nearly an order of magnitude lower than monolingual Hindi (38.18 ms) or isolated English (37.58 ms) baselines—thereby confirming the efficacy and necessity of customized modeling for code-switched speech.
📝 Abstract
Code-mixed speech poses unique challenges to forced alignment: expanded inventories, orthographic errors, and speaker variation. We evaluate forced alignment of Hindi-English code-mixed speech using the Montreal Forced Aligner. We address 2 problems: (1) free variation involving native vs non-native pairs and (2) phonemic boundary detection for mid-utterance English words. Bootstrapping strategies substantially outperform unmodified lexicons. Acoustic models trained on sentence-level code-mixed data achieve a mean error of 4.15ms, ie. ten times lower than monolingual Hindi (38.18ms) or isolated English (37.58ms) alternatives. Principled lexicon design and code-mixed training data are both essential for reliable alignment of bilingual speech.