Corrective Forcing: Unified Post-Training for Diffusions and Flows in Generative Speech Enhancement

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决生成语音增强中扩散和流模型的训练-推理不匹配问题,引入了矫正强迫(CoF)后训练范式,通过自生成展开状态学习并纠正预测。
📝 Abstract
Diffusion and flow models, as promising generative paradigms for speech enhancement, face a training--inference mismatch: training uses analytical path states, whereas inference recursively evaluates models on self-generated rollout states along discretized sampling trajectories. This mismatch causes prediction and discretization errors to accumulate. To address it, we introduce Corrective Forcing (CoF), a post-training paradigm that forces diffusion and flow models to learn from self-generated rollouts and correct their predictions. CoF corrects clean-speech predictions on rollout states toward the ground truth under dynamic sampling schedules, exposing the model to varying inference conditions. It further regularizes local evolution using locally corrected counterfactual transitions as references for factual transitions. By expressing model outputs through a shared clean-speech prediction parameterization, CoF applies the same post-training objective across diffusion and flow formulations. Experiments with SB-VE and OT-CFM demonstrate improvements in perceptual quality and reconstruction fidelity, together with robust performance across different numbers of sampling steps.
Problem

Research questions and friction points this paper is trying to address.

Diffusion models
Flow models
Training-inference mismatch
Prediction errors
Discretization errors
Innovation

Methods, ideas, or system contributions that make the work stand out.

Corrective Forcing
post-training
self-generated rollouts
dynamic sampling schedules
local evolution regularization