EditVoice: Variable-Length Non-Autoregressive Zero-Shot TTS and Speech Editing with Edit Flows

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of non-autoregressive zero-shot text-to-speech (TTS) systems, specifically the requirement to pre-specify speech duration and the disjointed nature of editing and generation tasks. To this end, this work proposes Edit Flows, a framework that jointly updates speech content and duration to unify zero-shot synthesis and text-based editing. Methodologically, it introduces the first variable-length non-autoregressive model and designs a complementary prompt sampling strategy to enable generalized editing beyond the training distribution. The framework is further optimized through speech infilling training on large-scale GigaSpeech data. Experimental results demonstrate that the proposed approach achieves highly competitive performance on both the Seed-TTS and LibriSpeech benchmarks.
📝 Abstract
Recent non-autoregressive (NAR) zero-shot text-to-speech (TTS) models generate in parallel but typically require the target sequence length to be specified before generation. We introduce EditVoice, to our knowledge the first variable-length NAR zero-shot TTS model, which uses Edit Flows to jointly update speech content and sequence length through insertions, deletions, and substitutions. EditVoice adopts speech-infilling training, which unifies zero-shot TTS and text-based speech editing and allows both prefix and suffix speech prompt placements at inference. We introduce Complementary Prompt Sampling (CPS) to leverage the complementary Edit Flow predictions induced by the two prompt placements. We further find that EditVoice can edit source and model-generated speech beyond its training sources. We use this generalization for end-to-end editing and training-free post-generation refinement. With the Edit Flow model trained on 10K h of GigaSpeech, EditVoice demonstrates competitive zero-shot TTS performance on Seed-TTS Eval EN and LibriSpeech-PC and speech editing performance on RealEdit.
Problem

Research questions and friction points this paper is trying to address.

Non-Autoregressive TTS
Zero-Shot Text-to-Speech
Variable-Length Generation
Speech Editing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Non-Autoregressive TTS
Edit Flows
Speech Infilling
Complementary Prompt Sampling
Zero-Shot Speech Editing
H
Hongyao Deng
School of Informatics, Xiamen University, China
Wenhao Guan
Wenhao Guan
Xiamen University
speech
X
Xuetao Lin
School of Informatics, Xiamen University, China
P
Peijie Chen
School of Informatics, Xiamen University, China
Weijie Wu
Weijie Wu
Roblox
Computer Networks
L
Lin Li
School of Electronic Science and Engineering, Xiamen University, China
Q
Qingyang Hong
School of Informatics, Xiamen University, China