DiTAR+: Dual Optimization for Robust Autoregressive Diffusion Speech Synthesis

📅 2026-09-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决长语音合成不稳定问题,提出DiTAR+框架,通过扩展历史接收域和解耦语义对齐与声学细节重建来提高生成鲁棒性。
📝 Abstract
Continuous-latent Autoregressive Diffusion Transformer (AR-DiT) models have demonstrated immense potential in zero-shot speech generation. However, they still suffer from limited decoding stability when synthesizing long utterances or complex linguistic structures. This instability primarily stems from a restricted historical receptive field and an acoustic inertia dependency within the diffusion decoder, which causes the model to ignore semantic conditions. To address these challenges, we propose DiTAR+, a dual-optimization framework. First, we introduce Dilated Context Sampling to expand the macro-level historical receptive field without violating physical temporal continuity, thereby preventing cumulative error propagation. Second, we propose Hierarchical Acoustic Masking to prevent shallow layers from attending to acoustic pre-context, explicitly decoupling semantic alignment from acoustic detail reconstruction. Extensive experiments show that our framework effectively mitigates pronunciation errors and semantic hallucinations, enhances generation robustness on challenging sentences, and maintains exceptionally high speaker similarity throughout the entirety of long-form utterances. On the linguistically challenging ZH-Hard set, DiTAR+ reduces the word error rate from 12.478% to 9.893%, and on extended utterances of 25 to 35 seconds it improves speaker similarity from 0.741 to 0.759 while simultaneously lowering the word error rate from 2.778% to 2.173%, outperforming both discrete-token and pure flow-matching baselines.
Problem

Research questions and friction points this paper is trying to address.

Autoregressive Diffusion
Decoding Stability
Linguistic Structures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dilated Context Sampling
Hierarchical Acoustic Masking
Dual-Optimization Framework
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Ziyu Zhang
ASLP@NPU, Northwestern Polytechnical University, Xi’an, China
T
Tianlun Zuo
ASLP@NPU, Northwestern Polytechnical University, Xi’an, China
Hanzhao Li
Hanzhao Li
Audio, Speech and Language Processing Group (ASLP@NPU), School of Computer Science, Northwestern
Speech SynthesisSpontaneous SpeechSpeech Codec
H
Haoyu Zhang
ASLP@NPU, Northwestern Polytechnical University, Xi’an, China
Lei Xie
Lei Xie
Northwestern Polytechnical University
speech processingspeech recognitionspeech synthesismultimediaartificial intelligence