Training-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance of existing pronunciation transcription methods on costly annotations or limited information sources by proposing ST2P, a training-free framework. This approach integrates lexical and acoustic cues using frozen pretrained models, replacing beam search with left-to-right greedy search under text constraints for efficient rescoring without additional training. Experimental results demonstrate that ST2P reduces the character error rate (CER) to 0.04% on Japanese while achieving three times the inference speed of the baseline. Furthermore, across multiple languages including Spanish, it outperforms both open-source multimodal large language models and conventional state-of-the-art methods. These findings establish ST2P as an efficient and accurate new paradigm for text-to-speech data preparation.
📝 Abstract
Accurate and efficient pronunciation transcription is essential for preparing text-to-speech training data at scale. Existing approaches have different limitations: grapheme-to-pronunciation (G2P) and speech-to-pronunciation (S2P) methods each capture only partial information, using only text or only speech, while speech-and-text-to-pronunciation (ST2P) methods use both but require costly pronunciation-annotated data. To address this problem, we propose a training-free ST2P pipeline that integrates both lexical and acoustic information at inference time. Lexical resources and G2P tools generate text-constrained candidates, and a left-to-right greedy search selects the best one using whole-sequence negative log-likelihoods from frozen pretrained S2P models. On three Japanese corpora, our method reduces Character Error Rate (CER) from 0.60--1.40\% (text-only baseline) to 0.04--0.17\% with reference transcripts, and 0.64--1.58\% with ASR transcripts. It outperforms all baselines, including a trained ST2P model and commercial multimodal LLMs. Our greedy search method is 3--3.5$\times$ faster than beam search at similar CER, and the cascade is 2$\times$ faster than direct decoding ensuring the efficiency and accuracy. In Spanish, French, and preliminary English, it also surpasses four open multimodal LLMs and the best traditional methods.
Problem

Research questions and friction points this paper is trying to address.

pronunciation transcription
grapheme-to-pronunciation
speech-to-pronunciation
text-to-speech
training-free
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-Free
Pronunciation Transcription
Acoustic Rescoring
Greedy Search
Text-Constrained Candidates
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.