🤖 AI Summary
This study investigates whether the performance gains of mask-replace training in zero-shot text-to-speech (TTS) stem solely from self-correction capabilities during inference. To address this, we propose DeMaR, a model integrating mask-replace training with confidence-based sampling. By disabling inference-time revisions on the LibriTTS dataset, we demonstrate that the observed improvements primarily originate from noise context augmentation and replacement supervision signals during training, rather than relying on token revision mechanisms at inference. Experimental results indicate that, under identical conditions, DeMaR achieves significantly lower word error rates compared to both autoregressive and pure mask-diffusion baselines. Crucially, this advantage persists even when inference-time revisions are disabled. These findings offer new insights into the training dynamics of discrete diffusion models for TTS.
📝 Abstract
Unlike autoregressive models, discrete diffusion-based models for zero-shot text-to-speech generate speech tokens in parallel and can revisit earlier predictions. Mask-and-replace training extends mask-only training by randomly replacing some tokens, and its gains are commonly attributed to self-correction, the ability to revise previously generated tokens. However, exposure to randomly perturbed context during training may itself improve generation, raising the question of whether these gains require inference-time token revision. To investigate this question, we use DeMaR, which combines mask-and-replace training with confidence-ranked mask-only sampling while preserving the total training corruption probability. Trained from scratch on LibriTTS, DeMaR achieves lower word error rates (WER) than autoregressive and mask-only diffusion baselines using the same speech tokenizer. This advantage persists when each token remains unchanged after first being unmasked. Matched training conditions on two heterogeneous speech tokenizers show that both noisy-context augmentation and replacement supervision improve WER under this restriction. These findings demonstrate training-side benefits of replacement beyond enabling inference-time token revision.