SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of simultaneously supporting natural language instruction-driven and reference-free zero-shot speech synthesis across diverse scenarios by proposing SwanTale, a unified generative framework that efficiently integrates both tasks within a single system for the first time. Built upon the newly curated high-quality annotated dataset SwanData-Caption, SwanTale leverages multimodal audio representations via SwanVAE, engram conditioning, a Unified Mixture-of-Experts (MoE) architecture, and GRPO reward-conditioned post-training to enable fine-grained joint control over semantic content, speaker style, and acoustic environment. Experimental results demonstrate that the proposed method achieves state-of-the-art performance across multiple metrics in both zero-shot and instruction-based speech generation, attains the highest expressiveness ratings, and successfully enables joint synthesis of multi-speaker speech with complex environmental sound effects.
📝 Abstract
Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, support acoustic scenes with environments and audio effects, and later reuse the designed voices. Therefore, it is important to support multi-speaker speech and audio generation for both instruct and zero-shot tasks. The instruct task requires a caption of the environment, speaker styles, and fine-grained content, while the zero-shot task uses reference audio together with the same fine-grained content. We address these tasks from both the data and model sides. First, we propose SwanData-Caption, which cleans raw speech and audio data, adds targeted synthetic coverage, and annotates diverse and accurate multi-level captions. Then, we propose SwanTale, a multi-speaker expressive speech and audio generation model that supports both zero-shot and instruct tasks. We introduce SwanVAE to support high-quality multi-audio-modality generation. Then, we adopt reward-conditioned quality control and Engram conditioning, along with Unified MoE for multi-task and multi-audio-modality modeling. In addition, we use curriculum learning and GRPO post-training to let the model progressively learn and strengthen its capabilities. Experimental results show that SwanTale leads on multiple key zero-shot and instruct metrics, achieves the best expressiveness scores in both tasks, and supports complex instruct generation involving multi-speaker speech and audio. Demos can be found at https://swanaigc.github.io/\#swantale.
Problem

Research questions and friction points this paper is trying to address.

multi-speaker speech generation
audio generation
instruct task
zero-shot task
voice design
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-speaker speech generation
zero-shot audio synthesis
instruction-driven generation
Unified MoE
SwanVAE
🔎 Similar Papers
No similar papers found.