🤖 AI Summary
This study addresses the high latency of autoregressive text-to-speech (TTS) systems and the reliance of non-autoregressive approaches on transcribed text by proposing an efficient zero-shot voice cloning framework. Methodologically, it replaces autoregressive decoding with masked prediction to enable transcript-free synthesis. Furthermore, a training-free acoustic length estimation mechanism is introduced to support cross-lingual generation and the use of reference audio devoid of lexical content. By integrating a non-autoregressive architecture with ReFlow distillation, the framework significantly accelerates the flow matching-based renderer. Experimental results demonstrate that this work achieves a tenfold increase in generation speed for long utterances while preserving high-fidelity audio quality and maintaining compatibility across multilingual scenarios.
📝 Abstract
TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning, such as requiring transcripts of the reference speech during inference. We present Tacit-TTS, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2. Our model replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation. Across two English and two Mandarin datasets, Tacit-TTS achieves competitive zero-shot quality while generating speech over 10x faster than IndexTTS2 for utterances longer than 5 seconds. Its transcript-free conditioning further supports cross-lingual and non-lexical references. We validate this capability using references from eight other languages, infant babble, and synthetic gibberish, where transcript-dependent systems often degrade or fail due to unreliable ASR transcripts.