🤖 AI Summary
This study addresses the challenges of poor word-level stress controllability and insufficient communicative accuracy in text-to-speech (TTS) synthesis by proposing a reinforcement learning-based non-autoregressive TTS system. The core innovation lies in the first application of the Group Relative Policy Optimization (GRPO) algorithm to TTS duration prediction. By constructing a word-level stress localization reward mechanism, the proposed approach enables direct optimization of the duration predictor and achieves precise stress control. Experimental results demonstrate that this system attains state-of-the-art performance in both stress controllability and objective evaluation metrics, while significantly outperforming existing baseline models in subjective preference tests.
📝 Abstract
Generating controllable and human-like emphasis remains an open challenge in text-to-speech, even when explicit emphasis control signals are provided in the text input, limiting the communicative accuracy of synthetic speech in real-world applications. Reinforcement learning has recently shown promise for post-training TTS systems to align with human preference, yet existing methods have not been applied to word-level prosodic control. We present EmphTTS, a non-autoregressive TTS system that applies Group Relative Policy Optimization (GRPO) to the duration predictor with an emphasis localization reward, enabling direct optimization for word-level emphasis. Evaluations show that EmphTTS achieves the best emphasis controllability and performs the best in emphasis objective evaluation. In subjective preference tests, EmphTTS is significantly preferred over synthetic groundtruth and most baselines. Ablation studies show that GRPO improves emphasis realization beyond supervised-finetuning-based duration modeling and simple speaking-rate adjustment, while alleviating the mismatch between the independently trained duration predictor and TTS model.