VoiceWeaver: Staged Learning of Structured Controls for Expressive Speech and Sound-Event Generation

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of superimposing heterogeneous control attributes—such as emotion, prosody, and sound events—in speech generation, which often induces multi-attribute interference and catastrophic forgetting. To mitigate these issues, we propose a staged learning framework that employs a shared text-audio architecture with structured label prefixes for unified modeling. Specifically, embedding decorrelation and attribute dropout are introduced to alleviate conditional interference, while replay distillation is incorporated to preserve previously acquired capabilities. Experimental results demonstrate that the proposed approach successfully enables the joint generation of expressive speech and sound events. The model achieves an accuracy of 86% for Chinese emotion control, joint emotion-event accuracies of 70% and 66% in Chinese and English, respectively, and approximately 59% accuracy when jointly controlling all three attributes.
📝 Abstract
Adding expressive and environmental controls to a speech generator requires learning heterogeneous attributes without losing earlier capabilities. VoiceWeaver addresses this problem with structured label prefixes and staged emotion-tone-event training in a shared text-audio model. Replay-based distillation retains earlier predictions, while embedding decorrelation and attribute dropout regularize conditioning. Evaluation separates single-attribute correctness, joint emotion-event generation, and three-attribute correctness. Rounded to whole percentage points, emotion accuracy is 86% in Chinese and 83% in English. Chinese emotion-event joint accuracy is about 70%, versus 56% for mixed-task training and 48% for Ming-omni-tts; English joint accuracy is 66%. Three-attribute accuracy is 59% in Chinese and 55% in English on 500 samples each (57% pooled). Text error rates exceed those of the external TTS baselines. Audio samples are available at https://anonymous.4open.science/api/repo/VoiceWeaver1-4D88/file/index.html.
Problem

Research questions and friction points this paper is trying to address.

speech generation
expressive control
environmental control
heterogeneous attributes
catastrophic forgetting
Innovation

Methods, ideas, or system contributions that make the work stand out.

Staged Learning
Structured Label Prefixes
Replay-based Distillation
Embedding Decorrelation
Attribute Dropout
🔎 Similar Papers
No similar papers found.
X
Xiaosu Su
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
Yun Cao
Yun Cao
researcher, tencent
CVGANs
Y
Yiping Ni
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
X
Xiaowei Yi
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China