🤖 AI Summary
This study addresses the challenge of speech enhancement for real-recorded audio, where paired clean targets are unavailable and optimization typically relies on differentiable rewards. To overcome this limitation, we propose a post-training optimization framework leveraging weak supervision from transcribed text. Specifically, the method employs Reinforce Adjoint Matching to post-train generative models such as FlowSE, enabling direct optimization of non-differentiable metrics like Word Error Rate (WER) without requiring paired data or differentiable reward functions. Evaluated on the CHiME-4 dataset, the proposed approach achieves a substantial WER reduction of 5.08%, while subjective listening tests confirm no degradation in perceptual quality. This work establishes an efficient new paradigm for non-differentiable metric-driven speech enhancement.
📝 Abstract
We adapt Reinforce Adjoint Matching (RAM), a reward-based post-training method, to generative speech enhancement (SE). Starting from a pretrained SE model, RAM tilts the model's conditional distribution toward outputs with higher reward. During training, the current model generates enhanced speech on-policy, evaluates each generated endpoint with a potentially non-differentiable reward, and analytically re-noises the endpoint to construct inputs for a reward-guided regression objective. This enables post-training directly on real recordings using weak supervision, such as text transcripts, without requiring paired clean speech targets or reward gradients. We investigate word error rate (WER)-based post-training and whether recognition performance can be improved without compromising perceptual speech quality. Experiments on real CHiME-4 recordings reduce WER by 5.08 percentage points relative to pretrained FlowSE without reducing any of the reported non-intrusive speech quality metrics. A subjective listening test at the default reward scale finds no statistically significant preference between the post-trained and pretrained models.