🤖 AI Summary
This study addresses the challenges of decoupling and continuously modeling semantic and acoustic features in high-fidelity human-like speech generation by proposing a fully continuous dual-encoder architecture. The method decomposes speech into joint semantic-acoustic representations, which are planned by a causal autoregressive Transformer and subsequently synthesized into 48kHz audio via a local diffusion model. Generation quality is further optimized through the synergistic integration of flow matching, supervised fine-tuning, and DiffusionNFT reinforcement learning. Experimental results demonstrate that Seed-TTS reduces the word error rate to 2.51%, achieving a 14.9% relative improvement over the baseline, while attaining state-of-the-art performance on both the InstructTTSEval and MDVD-Eval benchmarks.
📝 Abstract
We present JoyAI-Voice~2.0, an end-to-end anthropomorphic speech generation model built upon a fully continuous, dual-encoder architecture. Raw speech is encoded into continuous latents and partitioned into patches. Each patch is decomposed by a semantic-acoustic dual encoder into a semantically purified representation and an acoustic representation, which are fused and jointly fed to a causal autoregressive Transformer for planning. The Transformer predicts the conditioning for the next patch, and a local diffusion Transformer renders its full latents for 48\,kHz synthesis. The model is trained with a joint flow-matching and stop-prediction objective, followed by supervised fine-tuning and reinforcement learning with DiffusionNFT to improve model performance. It achieves the lowest average word error rate of 2.51\% on Seed-TTS, a 14.9\% relative reduction over the strongest baseline, and state-of-the-art attribute fidelity on InstructTTSEval in both Chinese and English, leading on 5 of 10 perceptual dimensions with the highest overall score of 0.893 on MDVD-Eval.