🤖 AI Summary
This study addresses the pronunciation representation mismatch caused by kanji orthography in reinforcement learning for Japanese text-to-speech (TTS) synthesis. To resolve this, we propose a policy optimization framework leveraging kana-domain automatic speech recognition (ASR) rewards. Specifically, we introduce a kana-domain character error rate as the reward signal to eliminate orthographic ambiguity, integrating Group Relative Policy Optimization (GRPO), a kana-transcription ASR model, and KL divergence regularization during training. Experimental results demonstrate that the proposed method reduces target kanji mispronunciation rates by approximately 26%, significantly mitigates output elongation artifacts, and substantially decreases the number of optimization steps required. Crucially, these improvements are achieved while preserving synthesized speech quality and speaker similarity.
📝 Abstract
Character error rate (CER) computed by automatic speech recognition (ASR) is widely used as an intelligibility reward for reinforcement learning (RL) post-training of text-to-speech (TTS) systems. For Japanese, however, orthographic CER introduces a representation mismatch for pronunciation-oriented optimization: distinct kanji readings may collapse to the same orthographic representation, while equivalent pronunciations may admit different orthographic forms. We instead compute CER in the kana domain using a kana-transcribing ASR model and reference readings (Kana-CER). Under matched group relative policy optimization (GRPO) conditions, Kana-CER reduces target-kanji reading error by approximately 26% relative to the orthographic CER reward, while maintaining comparable orthographic CER and similar speaker similarity and objective speech quality. It also reaches its best validation performance in substantially fewer optimization steps (4k vs. 18k). We further observe severe output elongation under unregularized Kana-CER optimization, which is substantially suppressed by KL regularization.