Graph Connectionist Temporal Classification for Phoneme Recognition

📅 2025-09-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Automatic phoneme recognition (APR) commonly relies on pseudo-phoneme labels generated by grapheme-to-phoneme (G2P) systems; however, standard Connectionist Temporal Classification (CTC) loss cannot model the inherent multi-pronunciation ambiguity in G2P outputs, resulting in poor robustness to label noise. To address this, we propose the first application of Graph-based Temporal Classification (GTC) to APR, wherein a phoneme sequence graph—explicitly encoding multiple pronunciation paths—serves as the supervision signal, enabling direct modeling of pronunciation uncertainty within the loss function. Our method supports end-to-end training without requiring post-processing or external pronunciation dictionaries. Evaluated on English and Dutch benchmark datasets, it achieves significant reductions in phoneme error rate, demonstrating both the effectiveness of integrating multi-pronunciation priors via GTC and its strong cross-lingual generalization capability.

Technology Category

Natural Language Processing: Language Grounding & Multi-modal NLPMachine Learning: Graph-based Machine LearningReasoning under Uncertainty: Graphical Models

Application Category

Graph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web search
📝 Abstract
Automatic Phoneme Recognition (APR) systems are often trained using pseudo phoneme-level annotations generated from text through Grapheme-to-Phoneme (G2P) systems. These G2P systems frequently output multiple possible pronunciations per word, but the standard Connectionist Temporal Classification (CTC) loss cannot account for such ambiguity during training. In this work, we adapt Graph Temporal Classification (GTC) to the APR setting. GTC enables training from a graph of alternative phoneme sequences, allowing the model to consider multiple pronunciations per word as valid supervision. Our experiments on English and Dutch data sets show that incorporating multiple pronunciations per word into the training loss consistently improves phoneme error rates compared to a baseline trained with CTC. These results suggest that integrating pronunciation variation into the loss function is a promising strategy for training APR systems from noisy G2P-based supervision.
Problem

Research questions and friction points this paper is trying to address.

Addressing ambiguity in phoneme-level annotations from G2P systems
Adapting Graph Temporal Classification for phoneme recognition training
Improving phoneme error rates by incorporating multiple pronunciations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adapts Graph Temporal Classification for APR
Trains from graph of alternative phoneme sequences
Integrates pronunciation variation into loss function
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Henry Grafé
Department of Electrical Engineering-ESAT, KU Leuven, Belgium
H
Hugo Van hamme
Department of Electrical Engineering-ESAT, KU Leuven, Belgium