Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead

📅 2026-06-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high latency in cascaded large language model and text-to-speech (TTS) systems, which arises from TTS models requiring full contextual input. To mitigate this, the authors propose S5-TTS, a streaming TTS model based on the T5 architecture that enables word-by-word incremental synthesis with minimal lookahead. S5-TTS integrates streaming language modeling and monotonic alignment learning, enhanced by a lookahead causal masking mechanism and convolution-augmented attention. Furthermore, interleaved multi-source knowledge distillation is employed to improve speech naturalness. Experimental results demonstrate that S5-TTS achieves comparable audio quality and high speaker similarity to its full-context counterpart, T5-TTS, while substantially reducing end-to-end response latency, making it well-suited for zero-shot dialogue scenarios.
📝 Abstract
Streaming text-to-speech synthesis in cascaded LLM-TTS systems still faces latency challenges as most TTS models require full context before initiating generation. We present S5-TTS, a streaming variant of T5-TTS that enables low-latency, word-by-word incremental speech synthesis through encoder-decoder language modeling and monotonic alignment learning. S5-TTS begins generating speech immediately after receiving the first few words, substantially reducing end-to-end response latency. To maintain quality under limited lookahead, we introduce a lookahead-causal masking mechanism with Conv-based auxiliary attention that preserves intelligibility and speaker similarity, and employ interleaved multi-source distillation to further restore naturalness. Experiments show that S5-TTS achieves comparable quality to full-context T5-TTS, supports zero-shot synthesis with high speaker similarity, and significantly reduces end-to-end latency for practical conversational AI systems.
Problem

Research questions and friction points this paper is trying to address.

streaming TTS
latency
limited lookahead
incremental speech synthesis
cascaded LLM-TTS
Innovation

Methods, ideas, or system contributions that make the work stand out.

streaming TTS
limited lookahead
monotonic alignment
causal masking
multi-source distillation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Muyang Du
NVIDIA, China
J
Jason Roche
NVIDIA, USA
J
Junjie Lai
NVIDIA, China