Score
Designs and implements algorithms, neural modules, and signal-processing pipelines that improve the fidelity and perceptual quality of audio waveforms by operating in time and/or frequency domains, including time-domain enhancement, frequency-based filtering, multi-resolution spectral reconstruction, and postprocessing after separation models (e.g., Demucs). Builds and analyzes components that selectively restore or boost high-frequency content (high-frequency enhancement blocks / hfeblock), apply low-frequency-guided recovery, and mitigate over-smoothing and nonlinear production artifacts to raise SNR and preserve fine waveform detail.
This study addresses the significant performance degradation of deep neural networks in audio source separation when inference is conducted at a lower sampling rate than that used during training, a phenomenon primarily attributed to the loss of high-frequency components. The authors systematically investigate this degradation mechanism and propose and validate the hypothesis that the mere presence of high-frequency information is more critical than its precise representation. To this end, they introduce two novel resampling strategies: noise-kernel resampling, which injects Gaussian noise into the high-frequency bands, and trainable-kernel resampling, which learns an interpolation kernel. Evaluated across multiple state-of-the-art music source separation models, both methods effectively mitigate performance degradation, with noise-kernel resampling demonstrating consistent robustness across architectures and offering a simple yet practical solution.
To address severe mid-to-high-frequency distortion, poor low-frequency fidelity, and inconsistent reconstruction in multi-source mixtures (e.g., instruments and vocals) for high-sample-rate audio restoration, this paper proposes the first explicit band-wise sequence modeling architecture. It decomposes the spectrogram into frequency bands to disentangle time–frequency dynamics and jointly models intra-band and inter-band temporal dependencies. The method integrates high-resolution time–frequency representations, hierarchical attention mechanisms, and a lightweight temporal encoder within a GAN-based band-decomposition generative framework. Evaluated on MUSDB18-HQ and MoisesDB under low-bitrate conditions (16–64 kbps), it significantly outperforms baselines such as SR-GAN, achieving gains of +2.1 dB in PSNR, +3.7% in STOI, and +0.42 in PESQ. The approach balances computational efficiency with superior detail recovery and fidelity preservation in complex mixed audio scenarios.
This study addresses the challenge of real-time music enhancement under strict causality and low-latency constraints, where music signals in recordings and live streams are commonly degraded by noise, reverberation, and spectral imbalances—conditions poorly handled by existing speech enhancement methods. The work proposes the first benchmark framework specifically designed for real-time music enhancement, integrating degradation-aware modeling, stereo processing, identity-preserving correction, and multidimensional evaluation. Leveraging compact causal neural networks, the authors systematically evaluate various models, including speech baselines, external music denoising systems, offline references, and customized MusicFilterNet-MS variants. Experiments demonstrate that all causal models operate faster than real time; however, enhancement efficacy is highly sensitive to degradation type, dataset, and evaluation metric, with improper processing often degrading audio quality rather than improving it.
This work addresses the problem of audio inpainting in the time-frequency domain, specifically targeting missing spectrogram columns. The authors propose an optimization-based approach that leverages a phase-aware prior by incorporating instantaneous frequency estimates to construct a reconstruction model that jointly exploits phase structure and signal priors. The resulting optimization problem is efficiently solved using a generalized Chambolle–Pock algorithm. The method achieves high reconstruction quality while significantly reducing computational cost, outperforming both state-of-the-art deep-prior neural networks and the Janssen-TF autoregressive approach in both objective metrics and subjective listening tests, thereby offering a favorable balance between performance and computational efficiency.
Audio denoising is critical for enhancing intelligibility and fidelity of complex audio signals such as music, yet discriminative models (e.g., U-Net) suffer from limitations in generation quality and fine-grained reconstruction. This paper introduces the first generative music denoising framework built upon the neural Audio Codec (DAC), achieved by adapting the Descript Audio Codec architecture into an end-to-end differentiable denoising system. We propose a multi-objective loss function jointly optimizing time-domain fidelity, spectral consistency, and perceptual quality. The model is trained on a large-scale, custom-synthesized noisy–clean paired dataset. Experiments demonstrate that our method significantly outperforms state-of-the-art discriminative and generative baselines across objective metrics (STOI, ESTOI, PESQ) and subjective listening tests. To our knowledge, this is the first work to validate the effectiveness and superiority of neural audio codecs for high-fidelity generative audio restoration.
This work proposes an end-to-end time-domain audio processing framework based on reservoir computing, addressing the limitations of traditional methods that rely on computationally intensive time–frequency transforms such as MFCCs and struggle to balance real-time performance, energy efficiency, and alignment with the human auditory system’s efficacy. By integrating biologically inspired auditory feature extraction with reservoir computing and replacing conventional frequency-domain transformations with lightweight convolutional operations, the proposed approach significantly reduces computational overhead while preserving discriminative feature representation. It eliminates the need for complex preprocessing and enables efficient, low-power real-time speech analysis, making it well-suited for embedded systems and voice-driven applications. This study thus establishes a highly energy-efficient and deployable paradigm for neuromorphic audio processing.
Bandwidth extension (BWE) aims to reconstruct high-frequency components from low-pass audio, a classic yet challenging audio generation task. This paper proposes a decoupled neural codec framework tailored for generative modeling, which—uniquely—integrates harmonic–percussive source separation (HPSS) into an end-to-end audio codec to explicitly disentangle and jointly optimize harmonic and percussive features. The method combines discrete audio tokenization, a Transformer-based language model for autoregressive high-frequency token prediction, and a joint training strategy that enhances representation learning and reconstruction consistency. Evaluated on multiple benchmark datasets, our approach achieves significant improvements over state-of-the-art methods in both objective metrics (PESQ, STOI) and subjective listening tests (MOS), demonstrating superior performance in high-fidelity high-frequency reconstruction and perceptual quality preservation.
This work addresses the challenge of high-fidelity, fine-grained controllable bandwidth extension for music audio degraded by limited bandwidth in historical archives—a task where existing generative models struggle to achieve both quality and precise control. The authors propose a single-step controllable bandwidth extension method based on Flow Matching, introducing for the first time a dynamic spectral contour (DSC) as a fine-grained conditioning signal. By integrating classifier-free guidance with DSC, the model enables accurate audio restoration with enhanced fidelity and controllability. Experimental results demonstrate that the proposed approach significantly outperforms prior methods in both perceptual quality and controllability, with DSC effectively facilitating high-precision conditional generation and establishing state-of-the-art performance in bandwidth extension tasks.
This work addresses the challenge of general speech restoration, which requires effectively modeling the complex structure of speech under diverse distortion conditions—a task at which existing methods struggle due to their inability to jointly capture spectral periodicity and multi-resolution frequency characteristics. To overcome this limitation, the paper proposes a novel state space model that, for the first time, incorporates spectral periodicity and multi-resolution analysis as inductive biases into speech restoration. The architecture features a frequency-domain GLP feature extraction module, a multi-resolution parallel time-frequency dual-processing structure, and a learnable mapping mechanism to efficiently integrate global, local, and periodic spectral patterns. The proposed method achieves state-of-the-art performance across multiple benchmarks while maintaining high computational efficiency.
Neural audio codecs often rely on data- or task-specific priors to disentangle frequency-band features, resulting in poor interpretability and limited generalizability. Method: We propose a generic soft disentanglement representation learning framework. It first applies spectral decomposition to project time-domain audio into orthogonal frequency-band subspaces; then employs a multi-branch encoder to model each band independently, jointly optimized via reconstruction and perceptual losses. Crucially, the framework imposes no assumptions about task structure or data distribution, enabling task-agnostic, soft intra-band semantic disentanglement. Contribution/Results: Experiments demonstrate significant improvements over state-of-the-art codecs in objective audio quality metrics (e.g., PESQ, STOI) and perceptual fidelity. Moreover, the learned representations exhibit strong generalization to downstream tasks—such as audio inpainting—and provide interpretable, structured frequency-band semantics without architectural or prior constraints.