🤖 AI Summary
This work addresses the high latency in cascaded large language model and text-to-speech (TTS) systems, which arises from TTS models requiring full contextual input. To mitigate this, the authors propose S5-TTS, a streaming TTS model based on the T5 architecture that enables word-by-word incremental synthesis with minimal lookahead. S5-TTS integrates streaming language modeling and monotonic alignment learning, enhanced by a lookahead causal masking mechanism and convolution-augmented attention. Furthermore, interleaved multi-source knowledge distillation is employed to improve speech naturalness. Experimental results demonstrate that S5-TTS achieves comparable audio quality and high speaker similarity to its full-context counterpart, T5-TTS, while substantially reducing end-to-end response latency, making it well-suited for zero-shot dialogue scenarios.
📝 Abstract
Streaming text-to-speech synthesis in cascaded LLM-TTS systems still faces latency challenges as most TTS models require full context before initiating generation. We present S5-TTS, a streaming variant of T5-TTS that enables low-latency, word-by-word incremental speech synthesis through encoder-decoder language modeling and monotonic alignment learning. S5-TTS begins generating speech immediately after receiving the first few words, substantially reducing end-to-end response latency. To maintain quality under limited lookahead, we introduce a lookahead-causal masking mechanism with Conv-based auxiliary attention that preserves intelligibility and speaker similarity, and employ interleaved multi-source distillation to further restore naturalness. Experiments show that S5-TTS achieves comparable quality to full-context T5-TTS, supports zero-shot synthesis with high speaker similarity, and significantly reduces end-to-end latency for practical conversational AI systems.