Phoneme-Guided Initialization for LLM-based Speech Recognition

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant performance degradation of large speech models in low-resource scenarios by proposing a phoneme-guided initialization method. The approach integrates the advantages of cascaded pipelines into an end-to-end framework: the audio encoder and the large language model are first pre-trained separately on speech-to-phoneme (S2P) and phoneme-to-grapheme (P2G) tasks, respectively, followed by end-to-end fine-tuning to optimize the model initialization strategy. Experimental results demonstrate that, on multilingual low-resource datasets, the proposed method matches or surpasses both cascaded baselines and end-to-end models without such initialization, effectively enhancing low-resource speech recognition capabilities.
📝 Abstract
Speech large language models (speech LLMs) perform well on automatic speech recognition (ASR) when sufficient paired speech-text data is available, but their performance degrades in low-resource settings. A cascaded pipeline that performs speech-to-phoneme (S2P) conversion followed by phoneme-to-grapheme (P2G) conversion has been shown to outperform end-to-end speech LLMs in this regime, suggesting that phoneme-mediated processing is beneficial when paired data is scarce. We propose \textit{phoneme-guided initialization}, a simple method that uses this insight within an end-to-end framework: we pre-train the audio encoder on S2P and the LLM on P2G tasks, then connect them and fine-tune the full model end-to-end on the target ASR task. Experiments on Japanese (CSJ), Chinese (AISHELL-1), and two low-resource languages from Common Voice 25.0 (Tatar and Urdu) show that our method matches or outperforms both the cascaded S2P-P2G baseline and the end-to-end model without P2G initialization.
Problem

Research questions and friction points this paper is trying to address.

Speech LLMs
Automatic Speech Recognition
Low-resource
Phoneme-guided initialization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Phoneme-Guided Initialization
Speech Large Language Models
Low-Resource ASR
Speech-to-Phoneme
Phoneme-to-Grapheme
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ryo Magoshi
Graduate School of Informatics, Kyoto University, Japan
S
Shinsuke Sakai
Graduate School of Informatics, Kyoto University, Japan
Tatsuya Kawahara
Tatsuya Kawahara
Professor, School of Informatics, Kyoto University
Speech Processingspeech recognitionNatural Language Processingdialogue