JoyAI-Voice 2.0: A Full-Continuous Autoregressive Speech Generation Model with Semantic-Acoustic Joint Representation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of decoupling and continuously modeling semantic and acoustic features in high-fidelity human-like speech generation by proposing a fully continuous dual-encoder architecture. The method decomposes speech into joint semantic-acoustic representations, which are planned by a causal autoregressive Transformer and subsequently synthesized into 48kHz audio via a local diffusion model. Generation quality is further optimized through the synergistic integration of flow matching, supervised fine-tuning, and DiffusionNFT reinforcement learning. Experimental results demonstrate that Seed-TTS reduces the word error rate to 2.51%, achieving a 14.9% relative improvement over the baseline, while attaining state-of-the-art performance on both the InstructTTSEval and MDVD-Eval benchmarks.
📝 Abstract
We present JoyAI-Voice~2.0, an end-to-end anthropomorphic speech generation model built upon a fully continuous, dual-encoder architecture. Raw speech is encoded into continuous latents and partitioned into patches. Each patch is decomposed by a semantic-acoustic dual encoder into a semantically purified representation and an acoustic representation, which are fused and jointly fed to a causal autoregressive Transformer for planning. The Transformer predicts the conditioning for the next patch, and a local diffusion Transformer renders its full latents for 48\,kHz synthesis. The model is trained with a joint flow-matching and stop-prediction objective, followed by supervised fine-tuning and reinforcement learning with DiffusionNFT to improve model performance. It achieves the lowest average word error rate of 2.51\% on Seed-TTS, a 14.9\% relative reduction over the strongest baseline, and state-of-the-art attribute fidelity on InstructTTSEval in both Chinese and English, leading on 5 of 10 perceptual dimensions with the highest overall score of 0.893 on MDVD-Eval.
Problem

Research questions and friction points this paper is trying to address.

speech generation
end-to-end model
anthropomorphic speech
attribute fidelity
word error rate
Innovation

Methods, ideas, or system contributions that make the work stand out.

Continuous Autoregressive Model
Semantic-Acoustic Joint Representation
Dual-Encoder Architecture
Diffusion Transformer
DiffusionNFT
🔎 Similar Papers
No similar papers found.
Yafeng Chen
Yafeng Chen
University of Science and Technology of China
Large Audio Language ModelSpeech Signal ProcessingDeep Learning
B
Boya Dong
Joy Future Academy, JD.com
Yankun Huang
Yankun Huang
Arizona State University
H
Hao Li
Joy Future Academy, JD.com
Jingdong Li
Jingdong Li
Li Auto
Signal Processing / Speech Enhancement / Text-to-Speech
X
Xiangyu Liang
Joy Future Academy, JD.com
H
Hao Ni
Joy Future Academy, JD.com
W
Wenchao Wang
Joy Future Academy, JD.com
Y
Yuxuan Wang
Joy Future Academy, JD.com
Z
Zhangyu Xiao
Joy Future Academy, JD.com
W
Wei Deng
Joy Future Academy, JD.com
Nan Duan
Nan Duan
JD.Com (now) | StepFun | Microsoft Research
NLPArtificial General Intelligence
Y
Yu Gu
Joy Future Academy, JD.com
Wenhao Guan
Wenhao Guan
Xiamen University
speech
W
Weisheng Han
Joy Future Academy, JD.com
Y
Yabin Li
Joy Future Academy, JD.com
Y
Yuan Liu
Joy Future Academy, JD.com
J
Jiaxin Ye
Joy Future Academy, JD.com
F
Fan Yu
Joy Future Academy, JD.com
L
Lin Zhu
Joy Future Academy, JD.com