π€ AI Summary
This study addresses the issue of incoherent spectral evolution during conditional flow matching (CFM) inference for speech synthesis by proposing a training-free frequency band enhancement strategy. The method leverages the discrete wavelet transform to dynamically modulate mel-spectrogram subbands, optimizing acoustic dynamics through synchronized low-frequency suppression and high-frequency boosting. To our knowledge, this constitutes the first training-free spectral synchronization mechanism tailored for CFM, overcoming the limitations of general-purpose diffusion model approaches in text-to-speech systems. Experimental results demonstrate that the proposed strategy reduces the number of neural function evaluations to 26 and yields a 61% improvement in FrΓ©chet Audio Distance, while preserving high audio quality and speaker similarity.
π Abstract
Conditional Flow Matching (CFM) models for text-to-speech (TTS) suffer from incoherent frequency evolution during inference. While similar spectral imbalances are addressed in diffusion models for other domains, those generic solutions fail to generalize to the inherently uncoordinated acoustic dynamics of CFM. We demonstrate that this issue can be effectively mitigated by introducing a novel training-free frequency-selective boosting strategy. Using the Discrete Wavelet Transform (DWT), our method dynamically modulates mel-spectrogram sub-bands during ODE integration, synchronizing spectral development by penalizing aggressive low-frequency growth and boosting lagging high-frequency details. Validated across diverse architectures (Matcha-TTS, F5-TTS, IndicF5), our approach reduces the required Number of Function Evaluations (NFE) from 32 to 26 and improves Frechet Audio Distance (FAD) by up to 61%, all without compromising mean opinion scores, speaker similarity, and speech intelligibility.