Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high latency of autoregressive text-to-speech (TTS) systems and the reliance of non-autoregressive approaches on transcribed text by proposing an efficient zero-shot voice cloning framework. Methodologically, it replaces autoregressive decoding with masked prediction to enable transcript-free synthesis. Furthermore, a training-free acoustic length estimation mechanism is introduced to support cross-lingual generation and the use of reference audio devoid of lexical content. By integrating a non-autoregressive architecture with ReFlow distillation, the framework significantly accelerates the flow matching-based renderer. Experimental results demonstrate that this work achieves a tenfold increase in generation speed for long utterances while preserving high-fidelity audio quality and maintaining compatibility across multilingual scenarios.
📝 Abstract
TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning, such as requiring transcripts of the reference speech during inference. We present Tacit-TTS, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2. Our model replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation. Across two English and two Mandarin datasets, Tacit-TTS achieves competitive zero-shot quality while generating speech over 10x faster than IndexTTS2 for utterances longer than 5 seconds. Its transcript-free conditioning further supports cross-lingual and non-lexical references. We validate this capability using references from eight other languages, infant babble, and synthetic gibberish, where transcript-dependent systems often degrade or fail due to unreliable ASR transcripts.
Problem

Research questions and friction points this paper is trying to address.

zero-shot voice cloning
non-autoregressive generation
transcript-free
text-to-speech
inference latency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Masked Non-Autoregressive Generation
Transcript-Free Voice Cloning
ReFlow Distillation
Training-Free Acoustic Length Estimation
Zero-Shot TTS
🔎 Similar Papers
No similar papers found.