Distill Locally, Schedule Globally: Flow Maps for Few-Step Text-to-Speech

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of excessive inference steps and high distillation costs in flow-matching text-to-speech (TTS) models by proposing a Local Flow Map Distillation (LFMD) framework. The method employs Euler mapping to circumvent teacher trajectory integration and integrates a dynamic programming-based TD-DP sampling scheduler, enabling flexible adaptation across multiple numbers of function evaluations (NFEs) without external metrics. Furthermore, soft dynamic time warping (DTW) temporal alignment-aware self-distillation is introduced to optimize single-step generation. Evaluated on the Seed-TTS dataset, the 1-NFE student model achieves a word error rate (WER) of 1.80%, closely approximating the performance of the 32-NFE teacher. These results demonstrate that LFMD significantly enhances both speech synthesis quality and training efficiency under low-step configurations.
📝 Abstract
Flow-matching text-to-speech (TTS) models achieve high synthesis quality but require many neural function evaluations (NFEs) to integrate their generative trajectories. Recent few-step flow-map distillation approaches for TTS construct targets from numerically integrated teacher trajectories, creating a trade-off between target accuracy and training cost. We propose Local Flow-Map Distillation (LFMD), which adapts Eulerian Map Distillation to conditional TTS and avoids teacher trajectory integration during target construction. For inference, we derive a sampling schedule (TD-DP) from teacher dynamics and consistency of the learned maps, with a single cost graph supporting multiple NFE budgets without external audio-metric evaluation. Because scheduling offers no flexibility at one NFE, we refine this regime with alignment-aware temporal self-distillation using soft-DTW. Across Seed-TTS and LibriSpeech-PC, LFMD improves low-NFE synthesis over a matched integral-distillation baseline. On Seed-TTS, the refined student reaches 1.80% WER with 1-NFE, compared with 1.76% for its 32-NFE teacher.
Problem

Research questions and friction points this paper is trying to address.

Text-to-Speech
Flow Matching
Few-Step Distillation
Neural Function Evaluations
Sampling Schedule
Innovation

Methods, ideas, or system contributions that make the work stand out.

Flow-Map Distillation
Text-to-Speech
Few-Step Generation
Sampling Schedule
Temporal Self-Distillation
🔎 Similar Papers
No similar papers found.