Transcript-Supervised Post-Training of Generative Speech Enhancement on Real Recordings via Reinforce Adjoint Matching

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of speech enhancement for real-recorded audio, where paired clean targets are unavailable and optimization typically relies on differentiable rewards. To overcome this limitation, we propose a post-training optimization framework leveraging weak supervision from transcribed text. Specifically, the method employs Reinforce Adjoint Matching to post-train generative models such as FlowSE, enabling direct optimization of non-differentiable metrics like Word Error Rate (WER) without requiring paired data or differentiable reward functions. Evaluated on the CHiME-4 dataset, the proposed approach achieves a substantial WER reduction of 5.08%, while subjective listening tests confirm no degradation in perceptual quality. This work establishes an efficient new paradigm for non-differentiable metric-driven speech enhancement.
📝 Abstract
We adapt Reinforce Adjoint Matching (RAM), a reward-based post-training method, to generative speech enhancement (SE). Starting from a pretrained SE model, RAM tilts the model's conditional distribution toward outputs with higher reward. During training, the current model generates enhanced speech on-policy, evaluates each generated endpoint with a potentially non-differentiable reward, and analytically re-noises the endpoint to construct inputs for a reward-guided regression objective. This enables post-training directly on real recordings using weak supervision, such as text transcripts, without requiring paired clean speech targets or reward gradients. We investigate word error rate (WER)-based post-training and whether recognition performance can be improved without compromising perceptual speech quality. Experiments on real CHiME-4 recordings reduce WER by 5.08 percentage points relative to pretrained FlowSE without reducing any of the reported non-intrusive speech quality metrics. A subjective listening test at the default reward scale finds no statistically significant preference between the post-trained and pretrained models.
Problem

Research questions and friction points this paper is trying to address.

Generative Speech Enhancement
Real Recordings
Weak Supervision
Word Error Rate
Perceptual Speech Quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforce Adjoint Matching
Generative Speech Enhancement
Post-Training
Weak Supervision
Word Error Rate
🔎 Similar Papers
No similar papers found.