DriftTTS: Few-Step Text-to-Speech Without Distillation via Distribution-Matching Drift

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the generation efficiency limitations of existing few-step text-to-speech (TTS) systems, which typically rely on knowledge distillation or complex schedulers. We propose DriftTTS, introducing a novel teacher-free, non-adversarial distribution-matching drift mechanism that integrates a frozen masked autoencoder (MAE) feature space with online policy rollout techniques to enable efficient mel-spectrogram generation. This approach achieves high-quality few-step synthesis without requiring distillation. Evaluated on the LJSpeech dataset, DriftTTS attains a Mel-Cepstral Distortion (MCD) of 3.87 dB and a Word Error Rate (WER) of 3.7% using only four inference steps, alongside a subjective blind Mean Opinion Score (MOS) of 4.18. These results demonstrate performance comparable to Matcha-TTS, establishing DriftTTS as an efficient new paradigm for lightweight speech synthesis.
📝 Abstract
Few-step neural text-to-speech models often rely on short- ened diffusion or flow-matching schedules, or on distillation from pretrained multi-step teachers. To avoid these depen- dencies, we present DriftTTS, a few-step mel-spectrogram generator trained without a generative teacher, distillation, or adversarial discrimination. DriftTTS uses a distribution- matching drift objective in a mel-domain feature space defined by raw mels and a frozen masked-autoencoder encoder pretrained on the same LJSpeech training split. On-policy rollout trains the decoder on its own interme- diate states and supports inference up to the trained roll- out depth. On LJSpeech, DriftTTS at NFE=4 achieves 3.87 dB MCD and 3.7% WER, compared with 3.85 dB and 3.4% for Matcha-TTS. In a fully paired blind listen- ing test, DriftTTS obtains 4.18 MOS, compared with 3.96 for Matcha-TTS and 4.22 for ground truth. These results demonstrate competitive few-step synthesis without a pre- trained generative teacher. Code can be found at https: //github.com/BASHLab/driftTTS.git
Problem

Research questions and friction points this paper is trying to address.

Few-step Text-to-Speech
Distillation-free
Distribution-matching drift
Mel-spectrogram generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Few-Step Text-to-Speech
Distribution-Matching Drift
On-Policy Rollout
Masked Autoencoder
Distillation-Free
🔎 Similar Papers
No similar papers found.