🤖 AI Summary
Automatic phoneme recognition (APR) commonly relies on pseudo-phoneme labels generated by grapheme-to-phoneme (G2P) systems; however, standard Connectionist Temporal Classification (CTC) loss cannot model the inherent multi-pronunciation ambiguity in G2P outputs, resulting in poor robustness to label noise. To address this, we propose the first application of Graph-based Temporal Classification (GTC) to APR, wherein a phoneme sequence graph—explicitly encoding multiple pronunciation paths—serves as the supervision signal, enabling direct modeling of pronunciation uncertainty within the loss function. Our method supports end-to-end training without requiring post-processing or external pronunciation dictionaries. Evaluated on English and Dutch benchmark datasets, it achieves significant reductions in phoneme error rate, demonstrating both the effectiveness of integrating multi-pronunciation priors via GTC and its strong cross-lingual generalization capability.
📝 Abstract
Automatic Phoneme Recognition (APR) systems are often trained using pseudo phoneme-level annotations generated from text through Grapheme-to-Phoneme (G2P) systems. These G2P systems frequently output multiple possible pronunciations per word, but the standard Connectionist Temporal Classification (CTC) loss cannot account for such ambiguity during training. In this work, we adapt Graph Temporal Classification (GTC) to the APR setting. GTC enables training from a graph of alternative phoneme sequences, allowing the model to consider multiple pronunciations per word as valid supervision. Our experiments on English and Dutch data sets show that incorporating multiple pronunciations per word into the training loss consistently improves phoneme error rates compared to a baseline trained with CTC. These results suggest that integrating pronunciation variation into the loss function is a promising strategy for training APR systems from noisy G2P-based supervision.