🤖 AI Summary
This study addresses the challenge of superimposing heterogeneous control attributes—such as emotion, prosody, and sound events—in speech generation, which often induces multi-attribute interference and catastrophic forgetting. To mitigate these issues, we propose a staged learning framework that employs a shared text-audio architecture with structured label prefixes for unified modeling. Specifically, embedding decorrelation and attribute dropout are introduced to alleviate conditional interference, while replay distillation is incorporated to preserve previously acquired capabilities. Experimental results demonstrate that the proposed approach successfully enables the joint generation of expressive speech and sound events. The model achieves an accuracy of 86% for Chinese emotion control, joint emotion-event accuracies of 70% and 66% in Chinese and English, respectively, and approximately 59% accuracy when jointly controlling all three attributes.
📝 Abstract
Adding expressive and environmental controls to a speech generator requires learning heterogeneous attributes without losing earlier capabilities. VoiceWeaver addresses this problem with structured label prefixes and staged emotion-tone-event training in a shared text-audio model. Replay-based distillation retains earlier predictions, while embedding decorrelation and attribute dropout regularize conditioning. Evaluation separates single-attribute correctness, joint emotion-event generation, and three-attribute correctness. Rounded to whole percentage points, emotion accuracy is 86% in Chinese and 83% in English. Chinese emotion-event joint accuracy is about 70%, versus 56% for mixed-task training and 48% for Ming-omni-tts; English joint accuracy is 66%. Three-attribute accuracy is 59% in Chinese and 55% in English on 500 samples each (57% pooled). Text error rates exceed those of the external TTS baselines. Audio samples are available at https://anonymous.4open.science/api/repo/VoiceWeaver1-4D88/file/index.html.