🤖 AI Summary
Addressing unsupervised image denoising under real-world noise without paired clean-noisy data, this paper proposes the Mask, Inpaint and Denoise (MID) framework. MID constructs self-supervised synthetic training pairs by randomly masking input noisy images and jointly optimizes denoising and inpainting tasks. Crucially, it introduces an “input sparsification-driven” iterative noise modeling mechanism: a closed-loop optimization alternately refines a Gaussian noise sampler and a denoiser, enabling automatic calibration of the noise distribution—entirely without clean reference images. This effectively bridges the domain gap between synthetic and realistic noise. Evaluated on multiple real-noise benchmarks, MID consistently outperforms existing unsupervised methods, achieving state-of-the-art performance. The results empirically validate that sparsity priors—induced via random masking—are both sufficient and effective for real-world denoising.
📝 Abstract
Supervised training for real-world denoising presents challenges due to the difficulty of collecting large datasets of paired noisy and clean images. Recent methods have attempted to address this by utilizing unpaired datasets of clean and noisy images. Some approaches leverage such unpaired data to train denoisers in a supervised manner by generating synthetic clean-noisy pairs. However, these methods often fall short due to the distribution gap between synthetic and real noisy images. To mitigate this issue, we propose a solution based on input sparsification, specifically using random input masking. Our method, which we refer to as Mask, Inpaint and Denoise (MID), trains a denoiser to simultaneously denoise and inpaint synthetic clean-noisy pairs. On one hand, input sparsification reduces the gap between synthetic and real noisy images. On the other hand, an inpainter trained in a supervised manner can still accurately reconstruct sparse inputs by predicting missing clean pixels using the remaining unmasked pixels. Our approach begins with a synthetic Gaussian noise sampler and iteratively refines it using a noise dataset derived from the denoiser's predictions. The noise dataset is created by subtracting predicted pseudo-clean images from real noisy images at each iteration. The core intuition is that improving the denoiser results in a more accurate noise dataset and, consequently, a better noise sampler. We validate our method through extensive experiments on real-world noisy image datasets, demonstrating competitive performance compared to existing unsupervised denoising methods.