Can LLM Agents Automate Reinforcement Learning for Text-to-Speech?

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the labor-intensive manual hyperparameter tuning in reinforcement learning (RL) post-training for zero-shot text-to-speech (TTS) by investigating whether large language model (LLM) agents can automate this process. We propose AgenticTTS-Forge, a collaborative framework enabling LLM agents to autonomously optimize RL strategies for CosyVoice2, incorporating multi-stage trajectory auditing and independent observer evaluation for validation. Results demonstrate that the agent successfully recovers and improves training recipes, halving failure cases through autonomous carrier switching. Furthermore, our analysis reveals that measurement rather than reasoning constitutes the primary bottleneck, identifying three major automation pitfalls including data leakage and reward inflation. Consequently, we advocate for designing full-channel contract frameworks to ensure reliable automation.
📝 Abstract
Although reinforcement learning (RL) post-training repairs the localized segmental errors of zero-shot text-to-speech (TTS), arriving at a working recipe still relies on tedious manual tuning, and whether LLM agents can take over this research pipeline is unclear. We investigate this question with AgenticTTS-Forge, a collaborative workflow that structures human guidance and agentic execution around a shared workspace, applied to CosyVoice2-0.5B. To measure what the agent automates, we audit its trajectory stage by stage against the published recipe. To measure what it exploits, we score its policies with held-out observers hidden from the agent. Our results show that the agent recovers an underspecified recipe, improves it, and, when gains stall, surveys the literature unprompted and pivots from the LM carrier to the flow carrier, halving Bad cases. However, its autonomy exposes three traps across the data, proxy, and algorithm axes: the held-out set leaks through a channel the contract never reads, a self-shaped reward inflates the proxy where it is scored, and separately tuned policies do not compose additively. These findings show that the binding constraint is measurement rather than reasoning, and can inform the design of harnesses whose contracts read every channel the agent does.
Problem

Research questions and friction points this paper is trying to address.

Text-to-Speech
Reinforcement Learning
LLM Agents
Automation
Post-training
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Agents
Reinforcement Learning
Text-to-Speech
Automated Machine Learning
Evaluation Methodology
🔎 Similar Papers
No similar papers found.