Phonemizing User-Generated Text: A Benchmark, Taxonomy, and Compositional Approach

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究解决了用户生成文本的发音问题,通过引入UGTPhon基准和提出结合规范形式证据的组合方法来改进图素到音素转换模型。
📝 Abstract
Text-to-speech systems increasingly process user-generated text (UGT) such as ppl and imo, whose pronunciation must be inferred from the canonical rather than surface form. We introduce UGTPhon, the first grapheme-to-phoneme (G2P) benchmark for UGT in English, Vietnamese, and Korean, together with an inference-grounded taxonomy for fine-grained diagnosis. Existing G2P models and frontier LLMs exhibit a systematic canonical-to-non-canonical performance gap, reaching up to 66.8 PER points. As a benchmark baseline, we propose a simple compositional G2P approach that incorporates canonical-form evidence through exact-match lookup and staged decoding. Across matched ByT5 and Qwen2.5-0.5B backbones, explicit canonical-form modeling consistently reduces non-canonical G2P errors. The 0.5B variant also performs competitively with much larger few-shot frontier LLMs, highlighting the benefit of explicitly modeling canonical-form inference for UGT phonemization.
Problem

Research questions and friction points this paper is trying to address.

Text-to-speech
User-Generated Text
Grapheme-to-Phoneme
Canonical Form
Performance Gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

UGTPhon
Grapheme-to-Phoneme (G2P)
Canonical Form Inference
Compositional G2P Approach
User-Generated Text
🔎 Similar Papers
No similar papers found.