🤖 AI Summary
This study addresses the challenge of recovering original instrument tracks from mixed, mastered, or degraded recordings by overcoming nonlinear production effects and transmission degradations. To transcend the limitations of conventional linear separation, this work proposes a modular three-stage paradigm termed “restoration–separation–re-restoration.” Specifically, the framework first employs a mixture restoration model to reverse signal degradations, subsequently utilizes a multi-source separator to extract eight target instrument tracks, and finally applies track-specific fine-tuned expert networks to eliminate residual artifacts. Evaluations on the MSR Challenge test set demonstrate that each stage yields significant, progressive improvements in audio restoration quality. Furthermore, the code and pretrained models are publicly released to facilitate future research in this domain.
📝 Abstract
Music Source Restoration (MSR) seeks to recover original, unprocessed instrument stems from mixed, mastered, and possibly degraded recordings. Unlike conventional source separation, which treats the mixture as a linear sum of clean sources, MSR must additionally invert nonlinear production effects, such as equalization and compression, and transmission-related degradations, such as codec artifacts. We propose a three-stage framework built around this distinction: (1) a mixture restoration model that addresses degradation before separation, (2) a single model that separates the restored mixture into eight target stems (vocals, guitars, keyboards, synthesizers, bass, drums, percussion, and orchestra), and (3) stem-specific restoration experts fine-tuned on the separator's own residual artifacts. Each stage improves restoration quality over the previous one on the MSR Challenge test set. We release code and models to support future research in MSR at https://github.com/theMoro/music_source_restoration.