🤖 AI Summary
This work addresses the problem of audio inpainting in the time-frequency domain, specifically targeting missing spectrogram columns. The authors propose an optimization-based approach that leverages a phase-aware prior by incorporating instantaneous frequency estimates to construct a reconstruction model that jointly exploits phase structure and signal priors. The resulting optimization problem is efficiently solved using a generalized Chambolle–Pock algorithm. The method achieves high reconstruction quality while significantly reducing computational cost, outperforming both state-of-the-art deep-prior neural networks and the Janssen-TF autoregressive approach in both objective metrics and subjective listening tests, thereby offering a favorable balance between performance and computational efficiency.
📝 Abstract
We address the problem of time-frequency audio inpainting, where the goal is to fill missing spectrogram portions with reliable information. Despite recent advances, existing approaches still face limitations in both reconstruction quality and computational efficiency. To bridge this gap, we propose a method that utilizes a phase-aware signal prior which exploits estimates of the instantaneous frequency. An optimization problem is formulated and solved using the generalized Chambolle-Pock algorithm. The proposed method is evaluated against other time-frequency inpainting methods, specifically a deep-prior audio inpainting neural network and the autoregression-based approach known as Janssen-TF. Our proposed approach surpassed these methods by a large margin in the objective evaluation as well as in the conducted subjective listening test, improving the state of the art. In addition, the reconstructions are obtained with a substantially reduced computational cost compared to alternative methods.