🤖 AI Summary
This study addresses the labor-intensive manual hyperparameter tuning in reinforcement learning (RL) post-training for zero-shot text-to-speech (TTS) by investigating whether large language model (LLM) agents can automate this process. We propose AgenticTTS-Forge, a collaborative framework enabling LLM agents to autonomously optimize RL strategies for CosyVoice2, incorporating multi-stage trajectory auditing and independent observer evaluation for validation. Results demonstrate that the agent successfully recovers and improves training recipes, halving failure cases through autonomous carrier switching. Furthermore, our analysis reveals that measurement rather than reasoning constitutes the primary bottleneck, identifying three major automation pitfalls including data leakage and reward inflation. Consequently, we advocate for designing full-channel contract frameworks to ensure reliable automation.
📝 Abstract
Although reinforcement learning (RL) post-training repairs the localized segmental errors of zero-shot text-to-speech (TTS), arriving at a working recipe still relies on tedious manual tuning, and whether LLM agents can take over this research pipeline is unclear. We investigate this question with AgenticTTS-Forge, a collaborative workflow that structures human guidance and agentic execution around a shared workspace, applied to CosyVoice2-0.5B. To measure what the agent automates, we audit its trajectory stage by stage against the published recipe. To measure what it exploits, we score its policies with held-out observers hidden from the agent. Our results show that the agent recovers an underspecified recipe, improves it, and, when gains stall, surveys the literature unprompted and pivots from the LM carrier to the flow carrier, halving Bad cases. However, its autonomy exposes three traps across the data, proxy, and algorithm axes: the held-out set leaks through a channel the contract never reads, a self-shaped reward inflates the proxy where it is scored, and separately tuned policies do not compose additively. These findings show that the binding constraint is measurement rather than reasoning, and can inform the design of harnesses whose contracts read every channel the agent does.