Phonologically Informed Tokenization for German Speech Recognition: A Cross-Domain Study

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether phonologically informed tokenization can serve as an effective target for end-to-end German speech recognition and evaluates its cross-domain generalization. Building upon the wav2vec 2.0 and CTC framework, and integrating the Knuth-Liang algorithm, Pyphen, and grapheme-to-phoneme conversion, it systematically compares character, BPE, and syllable- or phoneme-based tokenizers across standard, dialectal, and spontaneous speech. The findings reveal that tokenizer selection depends primarily on vocabulary budget rather than linguistic distribution, with the acoustic encoder governing the error topology. While all tokenizers achieve comparable performance in-domain, syllable-aware tokenization demonstrates significant advantages under constrained vocabulary budgets, particularly for dialectal and spontaneous speech scenarios.
📝 Abstract
German is a morphologically rich language whose syllable structure is exceptionally well-predicted by the Knuth--Liang hyphenation algorithm. We ask whether phonologically informed tokenization can serve as a competitive target for end-to-end speech recognition. We compare three tokenizer families on the Omnilingual ASR wav2vec 2.0 backbone fine-tuned with CTC: the pretrained multilingual character inventory, a data-driven Byte-Pair Encoding (BPE) over orthography, and phonologically informed units from Pyphen syllabification and grapheme-to-phoneme conversion. Across 40 fine-tunes, we evaluate on three German test sets spanning orthogonal shifts: in-domain read speech, dialectal spontaneous speech, and standard-German spontaneous speech. In-domain, all phonologically informed tokenizers match BPE and the multilingual character baseline on both WER and CER. Under domain shift the picture splits along vocabulary size rather than the linguistic axis of variation: at small vocabularies, syllable-aware tokenization improves on dialectal speech, where phonetic surface forms vary but syllable structure is preserved, and stays ahead on spontaneous speech, where new word-forms violate vocabulary closure. A phoneme-level confusion analysis further shows that all tokenizers commit the same canonical function-word errors, indicating that the acoustic encoder, not the tokenizer, dominates the error topology. Our findings suggest that tokenizer choice may depend on the vocabulary budget as much as on the distribution shift expected at deployment rather than reducing to a single universal optimum.
Problem

Research questions and friction points this paper is trying to address.

German speech recognition
phonologically informed tokenization
cross-domain
end-to-end ASR
vocabulary budget
Innovation

Methods, ideas, or system contributions that make the work stand out.

Phonologically Informed Tokenization
Cross-Domain Speech Recognition
Syllabification
wav2vec 2.0
Vocabulary Budget
🔎 Similar Papers
2024-06-21arXiv.orgCitations: 0