Tracing a Sparse Emotion-Control Circuit in LLM-Based Text-to-Speech

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear mechanism by which emotion is transferred from reference audio to synthesized speech in LLM-based text-to-speech (LLM-TTS) models. To investigate this, we employ a causal analysis framework integrating encoder-decoder trajectory scoring, late residual direction scoring, and activation patching interventions to precisely localize the sparse control circuits governing emotional expression within the model. Our findings reveal that merely 5% of critical components suffice to recover 74%–88% of the emotional effect, establishing a compact causal transmission pathway. Furthermore, we demonstrate that specific attention heads and MLP layers exert significant regulatory influence over acoustic features such as pitch and energy. These results provide an interpretable foundation for understanding the mechanisms underlying emotion generation in LLM-TTS systems.
📝 Abstract
LLM-based text-to-speech (TTS) models can generate emotionally expressive speech, but how reference emotion is routed through the model and realized in decoded speech remains unclear. We introduce two emotion-sensitive metrics for matched neutral and emotional syntheses---a codec trajectory score and a late residual direction score---and use them to score activation-patching interventions. Under controlled matched-reference conditions, this analysis identifies a sparse source-to-readout component-level circuit: 23--27 attention heads and MLPs per emotion, roughly 5% of the components considered, recover or suppress 74--88% of the late emotion-readout shift on held-out cases. The circuit combines a shared component backbone with emotion-specific components; cross-emotion activation swaps reduce the target readout in 47 of 48 cases. In decoded speech, the same intervention produces consistent changes in pitch, energy, and spectral brightness over 24 matched pairs per emotion. A readout-matched residual-direction baseline produces only 17--27% of the intervention's pitch effect, showing that internal readout movement alone does not explain the decoded acoustic changes. These results trace a compact causal route from reference-derived prefix information to emotion-relevant properties of generated speech.
Problem

Research questions and friction points this paper is trying to address.

text-to-speech
emotion control
large language models
mechanistic interpretability
causal circuit
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Emotion-Control Circuit
Activation Patching
Mechanistic Interpretability
LLM-based Text-to-Speech
Causal Intervention
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hongfei Du
Department of Computer Science, William & Mary, Williamsburg, USA
J
Jiacheng Shi
Department of Computer Science, William & Mary, Williamsburg, USA
Yanfu Zhang
Yanfu Zhang
William&Mary
Y
Ye Gao
Department of Computer Science, William & Mary, Williamsburg, USA