Listen, Critique, and Refine: RL-Based Self-Refinement for Instruction-Following Speech Synthesis

๐Ÿ“… 2026-09-21
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ไธบ่งฃๅ†ณๅคๆ‚่ฏญ้ŸณๅˆๆˆๆŒ‡ไปค็š„ๅฎž็Žฐ้—ฎ้ข˜๏ผŒ้€š่ฟ‡ๅผบๅŒ–ๅญฆไน ่ฎญ็ปƒๅคงๅž‹้Ÿณ้ข‘่ฏญ่จ€ๆจกๅž‹่ฟ›่กŒ่‡ชๆˆ‘ๅๆ€ๅ’Œไผ˜ๅŒ–๏ผŒๆๅ‡ๅˆๆˆ่ดจ้‡ใ€‚
๐Ÿ“ Abstract
Large Audio Language Models (LALMs) can follow diverse instructions to synthesize speech in specified styles. However, complex instructions that require simultaneous control over pitch dynamics, speaking rate, and emotional tone often exceed what a single-pass generation can faithfully realize. While recent reasoning models have shown that intermediate "thinking" tokens improve output quality, this paradigm has been confined to the text modality. In this work, we extend reasoning to the audio token space by training a LALM with reinforcement learning to reason over its own speech output. The model first generates a draft speech as a form of audio-token reasoning, critiques its own generation by reflecting on the acoustic realization in text, and then produces a refined version conditioned on both the first-pass speech and the critique, all within a single model. After RL training, the refined two-hop outputs achieve a relative improvement of 7.15\% on the InstructTTSEval benchmark, demonstrating the model's reflective ability.
Problem

Research questions and friction points this paper is trying to address.

Large Audio Language Models
complex instructions
speech synthesis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Self-Refinement
Audio Token Reasoning
Instruction-Following Speech Synthesis
๐Ÿ”Ž Similar Papers
No similar papers found.
C
Chee-En Yu
Graduate Institute of Electrical Engineering, National Taiwan University, Taiwan
Yi-Cheng Lin
Yi-Cheng Lin
National Taiwan University
Speech ProcessingMachine LearningFairness
Sung-Feng Huang
Sung-Feng Huang
Research Scientist, Nvidia
machine learningspeech processingnatural language processing
Y
Yun-Shao Tsai
Graduate Institute of Communication Engineering, National Taiwan University, Taiwan
H
Ho-Lam Chung
Graduate Institute of Communication Engineering, National Taiwan University, Taiwan
Xuanjun Chen
Xuanjun Chen
National Taiwan University
Speech ProcessingMachine LearningGenerative AIDeepfakes
Hung-yi Lee
Hung-yi Lee
National Taiwan University
deep learningspoken language understandingspeech processing