SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns

πŸ“… 2026-03-12
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge of general speech restoration, which requires effectively modeling the complex structure of speech under diverse distortion conditionsβ€”a task at which existing methods struggle due to their inability to jointly capture spectral periodicity and multi-resolution frequency characteristics. To overcome this limitation, the paper proposes a novel state space model that, for the first time, incorporates spectral periodicity and multi-resolution analysis as inductive biases into speech restoration. The architecture features a frequency-domain GLP feature extraction module, a multi-resolution parallel time-frequency dual-processing structure, and a learnable mapping mechanism to efficiently integrate global, local, and periodic spectral patterns. The proposed method achieves state-of-the-art performance across multiple benchmarks while maintaining high computational efficiency.

Technology Category

Natural Language Processing: SpeechMachine Learning: Multimodal LearningSearch and Optimization: Mixed Discrete/Continuous Search

Application Category

Graph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphsSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Robustness and generalizability of Web mining methods
πŸ“ Abstract
General speech restoration demands techniques that can interpret complex speech structures under various distortions. While State-Space Models like SEMamba have advanced the state-of-the-art in speech denoising, they are not inherently optimized for critical speech characteristics, such as spectral periodicity or multi-resolution frequency analysis. In this work, we introduce an architecture tailored to incorporate speech-specific features as inductive biases. In particular, we propose Frequency GLP, a frequency feature extraction block that effectively and efficiently leverages the properties of frequency bins. Then, we design a multi-resolution parallel time-frequency dual-processing block to capture diverse spectral patterns, and a learnable mapping to further enhance model performance. With all our ideas combined, the proposed SEMamba++ achieves the best performance among multiple baseline models while remaining computationally efficient.
Problem

Research questions and friction points this paper is trying to address.

speech restoration
spectral periodicity
multi-resolution frequency analysis
speech denoising
frequency bins
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speech Restoration
State-Space Models
Spectral Periodicity
Multi-resolution Analysis
Inductive Bias
Y
Yongjoon Lee
Korea Advanced Institute of Science and Technology (KAIST), Daejeon, Korea
J
Jung-Woo Choi
Korea Advanced Institute of Science and Technology (KAIST), Daejeon, Korea