Music Restoration via Latent Operator Optimization and Diffusion Model Priors

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenging problem of universal music signal restoration under unknown distortion types and without paired training data. The authors propose LOUDAR, a novel method that models distortion as an input-adaptive, learnable operator in the latent space of a pretrained audio autoencoder. By alternately optimizing clean latent variables and operator parameters—guided by an unconditional latent diffusion model as a prior—the approach enables unsupervised, general-purpose audio restoration. Experiments demonstrate that LOUDAR significantly outperforms degraded inputs across multiple tasks, including vocal effect removal, singing voice restoration, and guitar distortion elimination, achieving competitive or superior performance compared to existing supervised and unsupervised baselines in both waveform- and latent-space evaluation metrics.
📝 Abstract
Music restoration seeks to recover a clean signal from an observed recording degraded by an unknown effect, distortion, or corruption. Existing systems often rely on paired training data and distortion-specific supervision, which limits their use when the forward process is not known in advance. We propose LOUDAR (Latent-space Optimization of Unknown Distortion for Audio Restoration) a general-purpose restoration method that operates in the latent space of a pretrained audio autoencoder and models the unknown distortion as a learnable latent operator. At inference time, LOUDAR alternates between estimating the clean latent variable and updating the latent operator parameters. An unconditional latent diffusion model provides a prior over clean audio and regularizes this inference by steering the latent estimate toward the manifold of clean recordings. Because the degradation model is adapted per input, the approach is broadly applicable across diverse restoration problems. We evaluate LOUDAR on singing voice effect removal and restoration, as well as guitar distortion removal, and show that it consistently improves over degraded inputs and is competitive with supervised and unsupervised baselines in waveform and latent domains.
Problem

Research questions and friction points this paper is trying to address.

music restoration
unknown distortion
audio degradation
blind restoration
signal recovery
Innovation

Methods, ideas, or system contributions that make the work stand out.

latent operator optimization
diffusion model prior
audio restoration
unsupervised learning
latent space
M
Michal Švento
Dept. of Telecommunications, Brno University of Technology, Czech Republic
E
Eloi Moliner
Acoustic Lab, Dept. of Information and Communications Engineering, Aalto University, Finland
V
Valtteri Kallinen
Acoustic Lab, Dept. of Information and Communications Engineering, Aalto University, Finland
Lauri Juvela
Lauri Juvela
Assistant Professor, Machine Learning in Speech and Language Technology, Aalto University
generative deep learningspeech synthesismachine learning for audiospeech signal processing
Vesa Välimäki
Vesa Välimäki
Professor of Audio Signal Processing, Aalto University, Espoo, Finland
Audio Signal ProcessingAcoustic Signal ProcessingAudio EngineeringMusic TechnologySound and Music Computing
P
Pavel Rajmic
Dept. of Telecommunications, Brno University of Technology, Czech Republic