π€ AI Summary
This study addresses the coarse granularity, poor consistency, and high fine-tuning costs of emotion control during text-to-speech (TTS) inference by proposing a lightweight activation steering framework based on Qwen3-TTS. Specifically, steering vectors are injected into a frozen base model to enable precise emotional modulation. Furthermore, a two-pass generation pipeline incorporating multi-expert low-rank transformations and straight-through estimators is introduced to overcome discrete token limitations, facilitating continuous emotion intervention. Experimental results demonstrate that the proposed method improves emotion scores by up to 7.12 times and achieves a speaker identity retention rate 1.46 times higher than the baseline under strong preference settings, significantly outperforming existing approaches.
π Abstract
Pretrained text-to-speech (TTS) models can generate expressive speech, but reliable inference-time emotion control remains challenging: prompts and reference audio offer coarse, inconsistent control, whereas specialized conditioning and model adaptation require costly training. We present SteerSpeech, a lightweight activation-steering framework that controls emotion by injecting steering vectors into hidden activations. For each target emotion we train a lightweight low-rank transform, using a multi-expert objective that encourages monotonic emotion control while preserving speaker identity and linguistic content, constraining steering drift, and keeping the TTS backbone frozen. To optimize through discrete speech tokens, we introduce a two-pass generation-and-replay pipeline using a straight-through estimator to backpropagate expert supervision through sampled tokens. At inference, a target-emotion steering direction is optimized with its respective transform and injected into the base TTS model. Objective and subjective evaluations with Qwen3-TTS across seen, unseen, and accented speakers show stronger continuous emotion control with limited speaker and content degradation. SteerSpeech achieves 1.08x-7.12x baseline target-emotion scores and for a representative emotion subjectively, it receives 78.1%-96.8% intensity preference and 1.43x-1.46x speaker-identity preservation at high steering strengths.