ReaFlow-TTS: Realization-Conditioned Flow Matching for High-Quality and Controllable Speech Synthesis

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the marginalization of speech variation and the absence of semantically controllable interfaces in flow matching-based text-to-speech (TTS) systems by proposing a realization-conditioned flow matching framework. Methodologically, it introduces utterance-level stochastic latent variables to guide velocity prediction along generative trajectories, thereby preserving prosodic diversity while imposing valence-arousal-dominance (VAD) semantic constraints. By pioneering the integration of stochastic latents with VAD space, the framework enables direct, fine-grained attribute manipulation during inference without requiring target reference speech. Experimental results demonstrate that the proposed approach significantly enhances synthesis quality and effectively reproduces acoustic tendencies such as pitch and energy. Furthermore, subjective evaluations confirm that the method achieves precise, graded VAD control with negligible degradation in naturalness.
📝 Abstract
In flow-matching text-to-speech (TTS), different speech realizations can induce different target velocities under the same generation conditions. A deterministic velocity field trained with squared error predicts their conditional mean, thereby marginalizing realization-dependent variation. Meanwhile, modeling such variation does not inherently provide a semantically interpretable interface for attribute manipulation. We propose ReaFlow-TTS, a realization-conditioned flow-matching framework that introduces an utterance-level stochastic realization latent and uses it to condition velocity prediction throughout the generation trajectory. We further impose valence-arousal-dominance (VAD) semantics on the realization space, enabling direct and graded attribute manipulation without target speech at inference. Experiments demonstrate improved synthesis quality over a matched full-mask baseline and reproducible latent-induced pitch, energy, and timing tendencies across initial-noise samples, providing behavioral evidence that the latent is used as a reusable realization condition. Subjective evaluation further demonstrates graded VAD manipulation across generation contexts with only modest changes in naturalness.
Problem

Research questions and friction points this paper is trying to address.

Flow-matching TTS
Speech synthesis
Realization variation
Controllable attribute manipulation
VAD semantics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Flow Matching
Realization Latent
Controllable TTS
VAD Semantics
Speech Synthesis
🔎 Similar Papers
No similar papers found.