Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that text alignment fails to effectively predict downstream TTS acoustic similarity when natural language style descriptions serve as pseudo-labels. To overcome this, we propose a speech-rewarded style planning method that departs from conventional reliance on textual fidelity by directly employing target speech token likelihood as the reward signal. Specifically, we train a large language model-based style planner using the Group Relative Policy Optimization (GRPO) algorithm to achieve end-to-end voice style control through candidate instruction generation. Experimental results on the ISCSLP corpus demonstrate that the proposed approach significantly improves voice style and emotional similarity, reduces Mel-cepstral distortion, and enhances contextual appropriateness.
📝 Abstract
Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS). However, using descriptions as pseudo-labels compresses target acoustics into text, and descriptive fidelity need not imply effective control of a particular synthesizer. We empirically show that speech-text alignment only weakly predicts downstream acoustic similarity among candidate instructions for the same utterance. We therefore propose Speech-Rewarded Style Planning (SRSP), which trains a text-based style planner through a frozen downstream TTS model. Given dialogue history and response text, the planner generates candidate instructions and is optimized with group-relative policy optimization (GRPO), using the teacher-forced likelihood of target speech tokens as the reward. On an English subset of the ISCSLP 2026 CoT-TTS corpus, SRSP achieves higher speech-style and emotion similarity to target speech and lower mel-cepstral distortion than the Base LLM and target-audio-informed captioning baselines. LLM-based expressive speech evaluation further shows gains over all baselines in contextual appropriateness and reference consistency.
Problem

Research questions and friction points this paper is trying to address.

controllable text-to-speech
style planning
speech-text alignment
natural-language style descriptions
conversational TTS
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speech-Rewarded Style Planning
Group-Relative Policy Optimization
Conversational Text-to-Speech
Style Planning
Large Language Models
🔎 Similar Papers
No similar papers found.
S
Shiao Zhu
Institute of Science Tokyo, Tokyo, Japan
L
Lianbo Liu
Independent Researcher, Tokyo, Japan
S
Sizhen Lyu
Institute of Science Tokyo, Tokyo, Japan
Y
Yuzhe Wang
Institute of Science Tokyo, Tokyo, Japan
Sheng Li
Sheng Li
Institute of Science Tokyo (formerly TokyoTech) / RIKEN
Speech RecognitionDeep LearningNatural Language ProcessingMachine LearningEmbodied AI
T
Takahiro Shinozaki
Institute of Science Tokyo, Tokyo, Japan