Pronunciation-Oriented Reinforcement Learning for Japanese Text-to-Speech with Kana-Domain ASR Rewards

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the pronunciation representation mismatch caused by kanji orthography in reinforcement learning for Japanese text-to-speech (TTS) synthesis. To resolve this, we propose a policy optimization framework leveraging kana-domain automatic speech recognition (ASR) rewards. Specifically, we introduce a kana-domain character error rate as the reward signal to eliminate orthographic ambiguity, integrating Group Relative Policy Optimization (GRPO), a kana-transcription ASR model, and KL divergence regularization during training. Experimental results demonstrate that the proposed method reduces target kanji mispronunciation rates by approximately 26%, significantly mitigates output elongation artifacts, and substantially decreases the number of optimization steps required. Crucially, these improvements are achieved while preserving synthesized speech quality and speaker similarity.
📝 Abstract
Character error rate (CER) computed by automatic speech recognition (ASR) is widely used as an intelligibility reward for reinforcement learning (RL) post-training of text-to-speech (TTS) systems. For Japanese, however, orthographic CER introduces a representation mismatch for pronunciation-oriented optimization: distinct kanji readings may collapse to the same orthographic representation, while equivalent pronunciations may admit different orthographic forms. We instead compute CER in the kana domain using a kana-transcribing ASR model and reference readings (Kana-CER). Under matched group relative policy optimization (GRPO) conditions, Kana-CER reduces target-kanji reading error by approximately 26% relative to the orthographic CER reward, while maintaining comparable orthographic CER and similar speaker similarity and objective speech quality. It also reaches its best validation performance in substantially fewer optimization steps (4k vs. 18k). We further observe severe output elongation under unregularized Kana-CER optimization, which is substantially suppressed by KL regularization.
Problem

Research questions and friction points this paper is trying to address.

Japanese text-to-speech
character error rate
pronunciation optimization
reinforcement learning
representation mismatch
Innovation

Methods, ideas, or system contributions that make the work stand out.

Kana-CER
Reinforcement Learning
Japanese TTS
GRPO
KL Regularization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shiao Zhu
SB Intuitions Corp., Tokyo, Japan
L
Lianbo Liu
SB Intuitions Corp., Tokyo, Japan
K
Kai Washizaki
SB Intuitions Corp., Tokyo, Japan
K
Koki Nikaido
SB Intuitions Corp., Tokyo, Japan
Yui Sudo
Yui Sudo
SB Intuitions
Speech-to-SpeechSpeech recognitionRobot audition