🤖 AI Summary
This work addresses the challenging problem of blind estimation of unknown dynamic range compression (DRC) parameters and subsequent dry vocal recovery. The authors formulate DRC parameter estimation as a black-box optimization problem in a perceptual feature space, where parameters are inferred by minimizing the feature distance between the reconstructed signal and a reference signal in that space. Notably, the approach does not require a differentiable DRC model or a specific feature extractor, enabling compatibility with nonlinear transformations and histogram-based features, thereby overcoming limitations of conventional gradient-based methods. Experimental results demonstrate that the proposed method achieves state-of-the-art performance in both blind DRC parameter estimation and dry vocal restoration, yielding reconstruction quality that matches or surpasses existing best-performing approaches.
📝 Abstract
Dynamic Range Compression (DRC) is a widely used nonlinear audio effect whose parameters are often unknown, making blind estimation and inversion challenging. In this work, we formulate DRC parameter estimation as a black-box optimization problem in a perceptually motivated feature space. Given an observed signal and a reference representation, we estimate the parameters that minimize the distance between feature descriptors of the reconstructed and reference signals. Unlike gradient-based approaches, the proposed method does not require differentiability of the DRC model or the feature extraction pipeline, enabling the use of nonlinear and histogram-based descriptors. Experimental results demonstrate that the proposed method achieves competitive performance in blind parameter estimation and dry signal recovery, outperforming or matching state-of-the-art models in terms of reconstruction quality.