π€ AI Summary
This work addresses the challenge of general speech restoration, which requires effectively modeling the complex structure of speech under diverse distortion conditionsβa task at which existing methods struggle due to their inability to jointly capture spectral periodicity and multi-resolution frequency characteristics. To overcome this limitation, the paper proposes a novel state space model that, for the first time, incorporates spectral periodicity and multi-resolution analysis as inductive biases into speech restoration. The architecture features a frequency-domain GLP feature extraction module, a multi-resolution parallel time-frequency dual-processing structure, and a learnable mapping mechanism to efficiently integrate global, local, and periodic spectral patterns. The proposed method achieves state-of-the-art performance across multiple benchmarks while maintaining high computational efficiency.
π Abstract
General speech restoration demands techniques that can interpret complex speech structures under various distortions. While State-Space Models like SEMamba have advanced the state-of-the-art in speech denoising, they are not inherently optimized for critical speech characteristics, such as spectral periodicity or multi-resolution frequency analysis. In this work, we introduce an architecture tailored to incorporate speech-specific features as inductive biases. In particular, we propose Frequency GLP, a frequency feature extraction block that effectively and efficiently leverages the properties of frequency bins. Then, we design a multi-resolution parallel time-frequency dual-processing block to capture diverse spectral patterns, and a learnable mapping to further enhance model performance. With all our ideas combined, the proposed SEMamba++ achieves the best performance among multiple baseline models while remaining computationally efficient.