🤖 AI Summary
This study addresses the reliance of existing pronunciation transcription methods on costly annotations or limited information sources by proposing ST2P, a training-free framework. This approach integrates lexical and acoustic cues using frozen pretrained models, replacing beam search with left-to-right greedy search under text constraints for efficient rescoring without additional training. Experimental results demonstrate that ST2P reduces the character error rate (CER) to 0.04% on Japanese while achieving three times the inference speed of the baseline. Furthermore, across multiple languages including Spanish, it outperforms both open-source multimodal large language models and conventional state-of-the-art methods. These findings establish ST2P as an efficient and accurate new paradigm for text-to-speech data preparation.
📝 Abstract
Accurate and efficient pronunciation transcription is essential for preparing text-to-speech training data at scale. Existing approaches have different limitations: grapheme-to-pronunciation (G2P) and speech-to-pronunciation (S2P) methods each capture only partial information, using only text or only speech, while speech-and-text-to-pronunciation (ST2P) methods use both but require costly pronunciation-annotated data. To address this problem, we propose a training-free ST2P pipeline that integrates both lexical and acoustic information at inference time. Lexical resources and G2P tools generate text-constrained candidates, and a left-to-right greedy search selects the best one using whole-sequence negative log-likelihoods from frozen pretrained S2P models. On three Japanese corpora, our method reduces Character Error Rate (CER) from 0.60--1.40\% (text-only baseline) to 0.04--0.17\% with reference transcripts, and 0.64--1.58\% with ASR transcripts. It outperforms all baselines, including a trained ST2P model and commercial multimodal LLMs. Our greedy search method is 3--3.5$\times$ faster than beam search at similar CER, and the cascade is 2$\times$ faster than direct decoding ensuring the efficiency and accuracy. In Spanish, French, and preliminary English, it also surpasses four open multimodal LLMs and the best traditional methods.