🤖 AI Summary
This work addresses the challenge of zero-shot extension of pretrained single-step audio source separation models to multi-step inference. We propose a gradient-free iterative framework that, at each step, linearly interpolates the mixture with the previous step’s output and adaptively selects the interpolation weight via optimization to maximize a separation metric. We provide the first theoretical proof that such single-step models can be elevated to multi-step systems through this procedure, establishing an equivalence to denoising diffusion bridges. Under model smoothness constraints, we derive a robustness error bound for the separation metric. Empirically, our method consistently outperforms single-step baselines on speech enhancement and music separation tasks; the performance gain matches that achieved by scaling model capacity, increasing training data, or performing explicit multi-step joint training. Crucially, improvements generalize across all evaluation metrics—even those not explicitly optimized during inference.
📝 Abstract
Audio source separation aims to separate a mixture into target sources. Previous audio source separation systems usually conduct one-step inference, which does not fully explore the separation ability of models. In this work, we reveal that pretrained one-step audio source separation models can be leveraged for multi-step separation without additional training. We propose a simple yet effective inference method that iteratively applies separation by optimally blending the input mixture with the previous step's separation result. At each step, we determine the optimal blending ratio by maximizing a metric. We prove that our method always yield improvement over one-step inference, provide error bounds based on model smoothness and metric robustness, and provide theoretical analysis connecting our method to denoising along linear interpolation paths between noise and clean distributions, a property we link to denoising diffusion bridge models. Our approach effectively delivers improved separation performance as a"free lunch"from existing models. Our empirical results demonstrate that our multi-step separation approach consistently outperforms one-step inference across both speech enhancement and music source separation tasks, and can achieve scaling performance similar to training a larger model, using more data, or in some cases employing a multi-step training objective. These improvements appear not only on the optimization metric during multi-step inference, but also extend to nearly all non-optimized metrics (with one exception). We also discuss limitations of our approach and directions for future research.