A Training Criterion with Token-Level Tolerance to Transcription Ambiguity for Automatic Speech Recognition

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of conventional ASR training, which assumes unique transcriptions and overlooks local ambiguities. While existing Optimal Transducer Criterion (OTC) methods offer fault tolerance, they operate solely at the word level and risk discarding valid supervision. To overcome this, we propose a token-level OTC approach that refines wildcard arcs to subword granularity and integrates complementary word-level paths, thereby precisely preserving valid supervision signals. Furthermore, we introduce a predictive entropy-indexed scheduling mechanism that decouples length dependencies during training and dynamically localizes ambiguous regions. Extensive experiments demonstrate that the proposed method consistently outperforms CTC baselines across 25 tasks spanning 19 languages, achieving an average relative word error rate (WER) reduction of 9.45% and establishing state-of-the-art performance on all evaluated corpora.
📝 Abstract
Automatic speech recognition is typically trained assuming that the reference transcript is the only valid labeling of an utterance, yet even nominally verbatim transcripts contain localized differences in pronunciation, spelling, or lexical realization that the acoustics do not uniquely determine. Omni-temporal Classification (OTC) tolerates such noise by adding wildcard paths to the connectionist temporal classification (CTC) alignment graph, but its word-level arcs are too coarse, since bypassing one unsupported token discards supervision for the whole word. We move wildcard arcs to token granularity so unsupported tokens can be bypassed while the rest of the word stays supervised, and we combine token- and word-level arcs as complementary escape paths. Across 19 languages and three corpora, token-level OTC improves over CTC on all 25 tasks. We also replace epoch-indexed relaxation of the wildcard weights with a predictive-entropy-indexed schedule, which performs comparably while reducing dependence on training length. Combining this schedule with the hybrid graph gives the lowest mean word error rate (WER) on every corpus and a 9.45% average relative WER reduction over CTC. Independent validator transcriptions show that token-level models place significantly more wildcard-bypass probability than CTC on disputed characters, indicating that token-level tolerance targets localized transcript ambiguity.
Problem

Research questions and friction points this paper is trying to address.

Automatic Speech Recognition
Transcription Ambiguity
Omni-temporal Classification
Token-Level Tolerance
Connectionist Temporal Classification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Token-Level OTC
Transcription Ambiguity
Predictive Entropy Schedule
Hybrid Alignment Graph
Automatic Speech Recognition
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.