SE-MSB: End-to-End Unpaired Speech Enhancement using Mamba Schrödinger Bridges

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种使用Mamba Schrödinger Bridges的完全非配对语音增强框架,解决了未知声学特性环境下语音增强的问题,相比现有方法更高效且具有更好的泛化能力。
📝 Abstract
Speech enhancement (SE) models typically rely on supervised learning with paired data examples where clean speech is synthetically degraded. This paradigm limits performance in real-world scenarios where the target environment's specific acoustic characteristics are unknown. We propose a fully unpaired SE framework that uses principled Diffusion Schrödinger Bridges (DSB) to learn a stochastic transport process between a clean and a degraded speech distribution. Algorithms for learning transport maps are computationally heavy since they require simulating differential equations during training, usually at each training step. Therefore, we propose using a high-efficiency Mamba Diffusion Model designed for end-to-end waveform processing. We compare against state-of-the-art methods for speech enhancement, both paired and unpaired, as well as a classical signal processing algorithm. Experimental results show that we are on par or better than the baselines while being orders of magnitude faster during inference. Furthermore, we show that the flexibility of the DSB formulation allows our model to generalize across SE tasks, offering a robust and efficient solution for real-world speech restoration.
Problem

Research questions and friction points this paper is trying to address.

Speech Enhancement
Unpaired Data
Acoustic Characteristics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unpaired Speech Enhancement
Diffusion Schrödinger Bridges
Mamba Diffusion Model
End-to-End Waveform Processing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.