SCIC: Scope- and Codebook-Aware Instruction Conditioning for Speaker-Adapted Expressive TTS

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of clause-level prosody control and the limitations of global instructions on paragraph coherence in long-form livestreaming text-to-speech (TTS). We propose SCIC, a method enabling speaker-relative inline prosody control. SCIC introduces range-aware instruction conditioning based on residual vector quantization (RVQ) codebook properties, combined with temporal routing and tag-specific codebook weighting to optimize pitch and energy regulation within residual quantization. Furthermore, multi-reward Group Direct Preference Optimization (GDPO) post-training is incorporated to jointly enhance control precision and audio quality. Experimental results demonstrate that SCIC significantly improves pitch and energy control accuracy while generating more expressive hierarchical structures, maintaining low character error rates and high speaker similarity.
📝 Abstract
Long-form live-streaming TTS requires context-dependent prosody and paragraph-level coherence. However, many existing instruction-based TTS systems use global or uniform conditions, providing limited explicit control over clause-level relative prosodic changes. We introduce Speaker-Relative Inline Prosody Control, where each Pitch, Energy, or Speed instruction targets a clause relative to the preceding clause from the same speaker, while Pause uses an absolute duration interval. In codec-based TTS, Speed and Pause affect sequence length, whereas Pitch and Energy rely on residual codebooks. By analyzing Qwen3-TTS RVQ codebooks, we find that Energy concentrates in early residual codebooks, whereas Pitch accumulates across a deeper prefix. We therefore propose Scope- and Codebook-Aware Instruction Conditioning (SCIC), combining a Temporal Instruction Router for frame-level tag activation with Tag-Specific Codebook Weighting over residual codebooks. SCIC improves speaker-relative Pitch and Energy control over standard instruction fine-tuning using text-token tags. We further apply multi-reward GDPO post-training to jointly optimize control and quality, improving control accuracy while preserving CER and speaker similarity. In long-form synthesis, SCIC produces a more distinct paragraph-level expressive hierarchy than speaker-adapted SFT without instructions. Audio demos are available at: https://taoliveaigc.github.io/SCIC/
Problem

Research questions and friction points this paper is trying to address.

Expressive TTS
Prosody Control
Instruction Conditioning
Residual Codebooks
Speaker Adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Instruction Conditioning
Residual Vector Quantization
Codebook-Aware Weighting
Temporal Instruction Router
Multi-reward GDPO
🔎 Similar Papers
No similar papers found.
L
Longyu Lu
TaoLive-AIGC Team, Taobao & Tmall Group of Alibaba
Z
Zongwei Du
TaoLive-AIGC Team, Taobao & Tmall Group of Alibaba
M
Mengtao Xing
TaoLive-AIGC Team, Taobao & Tmall Group of Alibaba
Z
Zhuoqun Liu
TaoLive-AIGC Team, Taobao & Tmall Group of Alibaba
Z
Zifan Guan
TaoLive-AIGC Team, Taobao & Tmall Group of Alibaba
M
Meiguang Jin
TaoLive-AIGC Team, Taobao & Tmall Group of Alibaba
Junfeng Ma
Junfeng Ma
Mississippi State University
Design and ManufacturingLogisticsAI/MLHuman-Technology InteractionSustainability