Score
Design and build online text-to-speech systems that incrementally synthesize and emit audio as text (or other streaming inputs) arrive, operating without full buffering to minimize first-packet and end-to-end latency. Engineer lagged multi-track and incremental architectures that generate speech word-by-word or per-packet while preserving intelligibility and voice characteristics, including voice cloning and zero-shot speaker synthesis.
This work addresses the high latency in cascaded large language model and text-to-speech (TTS) systems, which arises from TTS models requiring full contextual input. To mitigate this, the authors propose S5-TTS, a streaming TTS model based on the T5 architecture that enables word-by-word incremental synthesis with minimal lookahead. S5-TTS integrates streaming language modeling and monotonic alignment learning, enhanced by a lookahead causal masking mechanism and convolution-augmented attention. Furthermore, interleaved multi-source knowledge distillation is employed to improve speech naturalness. Experimental results demonstrate that S5-TTS achieves comparable audio quality and high speaker similarity to its full-context counterpart, T5-TTS, while substantially reducing end-to-end response latency, making it well-suited for zero-shot dialogue scenarios.
Existing zero-shot streaming TTS systems predominantly rely on multi-stage discrete modeling, resulting in high latency, substantial computational overhead, and constrained speech quality. This paper proposes the first single-stage, zero-shot, low-latency streaming TTS framework, which innovatively employs interleaved autoregressive modeling over continuous mel-spectrograms: text tokens and acoustic frames are alternately fed into the model, integrated with streaming attention and a zero-shot speaker adaptation mechanism to enable end-to-end online speech synthesis. Evaluated on LibriSpeech, our method significantly outperforms existing streaming baselines; speech naturalness and speaker similarity approach those of offline systems, while end-to-end latency is reduced by over 50%. To the best of our knowledge, this is the first work achieving high-fidelity real-time synthesis without compromising strong cross-speaker generalization capability.
Existing zero-shot streaming TTS methods rely on future-text lookahead, incurring high latency. This paper proposes SMLLE, the first framework integrating real-time semantic modeling via Transducer with fully autoregressive streaming spectral reconstruction. It introduces semantic tokenization, duration-aligned learning, and a low-latency DeleteMechanism for dynamic frame-wise token deletion—enabling truly lookahead-free, high-fidelity frame-by-frame speech synthesis. SMLLE matches the naturalness of non-streaming sentence-level TTS while reducing end-to-end latency by over 40%, significantly enhancing practical streaming usability. Key contributions are: (1) a real-time co-architectural design that decouples semantic modeling from acoustic generation; (2) fine-grained temporal alignment and spectral reconstruction without future text; and (3) a controllable, stable, and computationally efficient streaming output scheduler via DeleteMechanism.
Zero-shot speech synthesis faces the challenge of simultaneously achieving high-fidelity audio quality, speaker consistency, and low latency—particularly in streaming systems. This work proposes an efficient, high-quality synthesis framework that integrates a causal variational autoencoder, block-wise autoregressive modeling, explicit speaker conditioning, and a bidirectional flow-matching head. A key innovation is guided-step distillation, which unifies classifier-free guidance with a multi-step solver into a single-step interval-conditioned student model, substantially reducing inference latency while preserving audio fidelity. Experimental results demonstrate competitive intelligibility and speaker similarity on LibriSpeech and Seed-TTS-Eval, with a 23.3% reduction in first-audio latency and a 40.8% improvement in real-time factor.
Traditional TTS systems require complete sentence input, resulting in high initial-word latency when cascaded with streaming LLMs—severely limiting real-time responsiveness in conversational AI. To address this, we propose the first end-to-end streaming TTS framework based on a decoder-only Transformer architecture. Our method introduces interleaved text–speech modeling and a next-token prediction loss, enabling truly incremental speech synthesis; it dynamically accepts text chunks and generates speech synchronously. Experiments demonstrate that our approach achieves state-of-the-art initial-word latency (<300 ms) while matching the naturalness of non-streaming TTS systems (MOS ≈ 4.2). This significantly improves both response efficiency and interaction fluency in cascaded LLM–TTS dialogue systems.
Existing text-to-speech (TTS) systems struggle to simultaneously achieve low latency and streaming capability due to autoregressive generation or multi-step flow matching. This work proposes FlashTTS, a natively streaming TTS framework that introduces a novel lagging multi-track architecture to eliminate sentence-level buffering. By integrating multi-token parallel prediction (MTP) with an X-pred mean-based flow-matching decoder, FlashTTS enables high-quality non-autoregressive acoustic generation in just two function evaluations. The method reduces first-packet latency to 325 ms—significantly outperforming strong baselines—while preserving excellent zero-shot voice cloning performance and cross-lingual intelligibility.
This study addresses the critical challenges of high latency and the trade-off between audio fidelity and real-time interruption in voice interaction models. We propose a low-latency, continuous streaming text-to-speech (TTS) model designed for intelligent agents, driven directly by LLM text streams. By innovatively incorporating control tokens and a dynamic silence generation mechanism, our approach enables mid-utterance interruptions without KV cache resets and supports silent output during idle states. The system maintains a modular architecture and high-fidelity audio quality while achieving persistent, always-on real-time responsiveness. Consequently, this work effectively resolves the fundamental difficulty in traditional architectures of balancing rapid response times with natural conversational interactivity.
This work addresses the unnatural inter-sentence silences—up to 9.6 seconds—in conventional real-time game commentary systems, which stem from strictly sequential text generation and speech synthesis. To mitigate this latency bottleneck, the authors propose an end-to-end low-latency commentary architecture that parallelizes large language model–based text generation with speech synthesis and incorporates a multi-candidate sentence pre-caching mechanism to enable immediate voice output at utterance boundaries. The proposed approach reduces the average inter-sentence silence duration to 0.3 seconds and improves temporal alignment with professional commentators by over 40%. A user study involving 120 experienced gamers demonstrates that the system significantly enhances the naturalness of speaking rhythm and overall immersion.
This work addresses the limitations of existing speech-to-text translation systems, which typically rely on cascaded architectures prone to error propagation, and current SpeechLLMs that fail to support truly real-time streaming translation, resulting in high latency. The paper proposes the first genuinely streaming SpeechLLM architecture, wherein a large language model learns an end-to-end mapping from speech to translated text and dynamically determines output timing—autonomously deciding when sufficient acoustic context is available to generate the next token, without relying on fixed intervals or waiting for complete utterances. By incorporating paralinguistic cues and training on automatically aligned speech–text data, the method achieves translation quality approaching that of non-streaming baselines across multiple languages while maintaining latency within 1–2 seconds, substantially enhancing practicality and responsiveness.
为解决低延迟语音对话系统中的流式文本到语音转换问题,提出X2Streaming-TTS框架,通过因果承诺和语音状态继承方法处理不确定前缀并保持声学连续性。