🤖 AI Summary
This study addresses the issue of content and speaker identity hallucinations frequently induced by generative speech enhancement. To mitigate this, we propose AuraSE, a novel framework that pioneers a dual-stream multimodal diffusion Transformer (MMDiT) flow-matching architecture to preserve the original acoustic pathway. Furthermore, it incorporates an online preference optimization (IPO) algorithm alongside multi-objective reward modeling to dynamically refine decoding strategies, enabling efficient fixed-step inference without utterance-level search. Experimental results demonstrate that AuraSE ranks first on 11 out of 12 metrics across synthetic test sets. Additionally, in real-world Deep Noise Suppression (DNS) blind evaluations, it achieves the highest DNSMOS scores and superior subjective listening quality ratings.
📝 Abstract
Generative speech enhancement models can produce cleaner and more natural-sounding speech than conventional discriminative approaches, but may hallucinate by changing speech content or speaker identity, even with transcript conditioning. We present AuraSE, a flow-matching framework that addresses hallucination through complementary modality and inference designs. First, a double-stream-to-single-stream multimodal Diffusion Transformer (MMDiT) allows transcript and acoustic representations to interact while preserving a dedicated pathway for the degraded input. Second, we find that the best decoder configuration, governed by guidance scale, sampling temperature, and step count, varies substantially across utterances. This observation motivates Inference Policy Optimization (IPO), an online, on-policy preference optimization method. IPO generates multiple candidates from the current model under different inference configurations, ranks them with a multi-objective reward, and learns from their relative preferences. AuraSE-IPO ranks first on 11 of 12 metrics across the synthetic test sets and obtains the highest DNSMOS and blind-listening scores among the evaluated systems on the real DNS blind test set. At deployment, it uses a fixed $10$-step ODE decoder without classifier-free guidance (CFG) or per-utterance configuration search.