Controlling Speaking Rate in Autoregressive TTS via Activation Steering

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of flexibly controlling speech rate in autoregressive text-to-speech (TTS) models after training. To this end, it proposes a retraining-free, inference-time speech rate control mechanism based on activation steering. Specifically, the method first learns a speed axis using synthetically time-scaled data. It then establishes a neutral operating point via single-decoder-block activation clamping and introduces intensity scaling. During inference, activations are projected and fixed to achieve precise rate regulation. The proposed approach enables stable and continuous speech rate adjustment without retraining, significantly outperforming conventional baselines while strictly preserving speaker identity and speech naturalness.
📝 Abstract
Autoregressive text-to-speech (TTS) systems synthesize natural speech but, once trained, offer little control over speaking rate. We show that speaking rate can be steered at inference time, without retraining, by clamping a single decoder block's activation along a discovered speed axis. A decoder-block analysis recovers the rate axis, a neutral operating point, and a per-step intensity scale; at inference, the activation's projection onto this axis is set to a fixed scalar. Learning this direction from synthetically time-stretched and time-compressed speech yields rate control that largely preserves speaker identity, generalizes across model architectures, and maintains high naturalness in objective and human evaluations. Unlike standard additive steering, which breaks at the slow extreme, clamping remains stable on all three systems tested; at moderate targets, the better rule depends on the model. Finally, we show that rate information is decodable across layers but causally steerable only within a mid-depth window, and demonstrate the effectiveness of our approach on the public Seed-TTS-Eval benchmark.
Problem

Research questions and friction points this paper is trying to address.

Autoregressive TTS
Speaking Rate Control
Activation Steering
Inference-time Control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Autoregressive TTS
Activation Steering
Speaking Rate Control
Inference-time Control
Clamping
🔎 Similar Papers