Prompt-Consistency Inference for Zero-Shot Flow-Matching Text-to-Speech Models

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the generation bias caused by prompt state drift during inference in zero-shot text-to-speech synthesis. To mitigate this issue, we propose a training-free prompt-consistent inference method built upon flow matching models and ordinary differential equation solvers. Specifically, this work introduces a novel closed-form solution-based prompt state correction mechanism that recovers prompt segments along the analytical trajectory to rectify the velocity field. Experimental results demonstrate that the proposed approach significantly improves both speaker similarity and intelligibility of synthesized speech without introducing additional network evaluation overhead. Consequently, this method offers an efficient and reliable training-free optimization paradigm for zero-shot speech synthesis.
📝 Abstract
In recent years, flow-matching models have produced significant improvements in zero-shot text-to-speech synthesis. Conditioned on an audio prompt and text, these models learn a velocity field and generate speech by iteratively solving an ODE. During inference, the solver evolves a single state spanning both the prompt and the region to be generated, although only the generated region is ultimately retained. The discarded prompt state, however, still matters; its intermediate values influence generation through the velocity field that couples the two regions. As sampling continues, this state can drift away from the prescribed conditional path, introducing a discrepancy into subsequent generation updates. Unlike the unknown generated trajectory, the prompt path is available in closed form from the reference audio and the initial noise. We exploit this observation with Prompt-Consistency Inference (PCI), a training-free rule that restores the prompt block to its analytic value before each velocity evaluation, while leaving the generated block unchanged. PCI improves speaker similarity and intelligibility across the evaluated flow-matching TTS backbones without additional network evaluations. Our ablation studies further show that PCI keeps post-step prompt discrepancies smaller and that corrections covering the later sampling stages recover much of the observed similarity gain.
Problem

Research questions and friction points this paper is trying to address.

zero-shot text-to-speech
flow-matching models
prompt drift
speaker similarity
inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Flow-Matching
Zero-Shot Text-to-Speech
Prompt-Consistency Inference
Training-Free
ODE Solver
🔎 Similar Papers
No similar papers found.
V
Vasily Zadorozhnyy
Applied Sciences Group, Microsoft Corporation, Redmond, USA
C
Can Goksen
Applied Sciences Group, Microsoft Corporation, Redmond, USA
Kazuhito Koishida
Kazuhito Koishida
Microsoft Corporation
Speech/Audio Recognition/Compression/Generation. Signal Processing and Machine Learning
D
Dung Tran
Applied Sciences Group, Microsoft Corporation, Redmond, USA