Score
Designs and implements image-denoising methods that reconstruct a clean image by fitting an untrained neural network (the deep image prior) to a single noisy observation, operating without externally paired training data. Analyzes and tunes the optimization, implicit or explicit regularization (e.g., sequential autoencoding-style constraints) and stopping criteria to suppress noise while preserving fine structures and improving signal-to-noise.
Addressing unsupervised image denoising under real-world noise without paired clean-noisy data, this paper proposes the Mask, Inpaint and Denoise (MID) framework. MID constructs self-supervised synthetic training pairs by randomly masking input noisy images and jointly optimizes denoising and inpainting tasks. Crucially, it introduces an “input sparsification-driven” iterative noise modeling mechanism: a closed-loop optimization alternately refines a Gaussian noise sampler and a denoiser, enabling automatic calibration of the noise distribution—entirely without clean reference images. This effectively bridges the domain gap between synthetic and realistic noise. Evaluated on multiple real-noise benchmarks, MID consistently outperforms existing unsupervised methods, achieving state-of-the-art performance. The results empirically validate that sparsity priors—induced via random masking—are both sufficient and effective for real-world denoising.
Existing learning-based image denoising methods rely on fixed noise priors, leading to poor generalization under varying real-world noise distributions. To address this, we propose a novel paradigm that decouples noise priors from image priors, and introduce a conditional optimization framework capable of estimating sensor-level noise priors directly from a single sRGB noisy image. Our key contributions are: (1) the first explicit modeling and incorporation of noise priors into denoising architectures; (2) a lightweight Local Noise Prior Estimator (LoNPE) network for pixel-wise noise prior estimation; and (3) a Conditional Denoising Transformer (CondFormer) that dynamically injects estimated noise priors into the denoising subspace via conditional self-attention. Extensive experiments demonstrate significant improvements over state-of-the-art methods on both synthetic and real-world datasets, with strong cross-device robustness and generalization capability. The code is publicly available.
Plug-and-Play (PnP) algorithms apply denoisers to progressively noise-decaying iterates, conflicting with diffusion models (DMs), which deploy denoisers exclusively on controllably noisy data. This inconsistency undermines theoretical alignment and practical performance. Method: We propose SNORE—a stochastic noise-level-adaptive regularization framework for image inverse problems (e.g., deblurring, inpainting). SNORE constructs a noise-level-matched stochastic gradient descent optimizer via explicit noise-aware regularization. Contribution/Results: This is the first PnP method to incorporate noise-perceptive stochastic regularization, unifying PnP and DM denoising logic while providing rigorous convergence and annealing-theoretic analysis. Experiments demonstrate that SNORE, when integrated with deep denoisers (e.g., DnCNN), achieves state-of-the-art performance on deblurring and inpainting—outperforming prior methods in PSNR, SSIM, and perceptual quality.
To address the limited performance of real-image denoising under extremely low-light conditions, this work introduces, for the first time, natural-language scene descriptions provided by photographers as explicit semantic priors into the denoising pipeline, moving beyond conventional purely data-driven paradigms. Methodologically, we propose a text-guided diffusion model that jointly leverages a CLIP text encoder and a U-Net-based denoising network to enable cross-modal conditional reconstruction. Evaluated on both synthetic and real low-light datasets, our approach achieves significant improvements in PSNR and SSIM; notably, in single-frame scenarios with extreme noise, structural and textural details are visibly restored. This work pioneers a language-prior-driven paradigm for real-image denoising, establishing an interpretable and controllable semantic guidance framework for low-light visual reconstruction.
This paper addresses the poor generalization of learning-based image restoration methods in real-world scenarios—a limitation stemming from significant domain shift between synthetic training data and real images. To bridge this gap, we propose a novel domain adaptation paradigm tailored to the noise space of diffusion models. Our key contributions are: (1) the first “denoising-as-adaptation” mechanism, which progressively aligns restoration outputs of synthetic and real images toward the clean distribution via multi-step conditional denoising and a domain-aligned diffusion loss; and (2) a channel-shuffling layer coupled with residual-swap contrastive learning to implicitly blur domain boundaries and suppress shortcut feature dependencies. Evaluated on denoising, deblurring, and deraining tasks, our method substantially outperforms existing domain-adaptive and blind restoration approaches, achieving state-of-the-art generalization performance on real-world images.
Real-world image denoising faces dual challenges: poor generalizability of handcrafted priors and the heavy reliance of deep learning methods on large-scale paired noisy-clean training data. To address this, we propose Net2Net—a novel framework that, for the first time, seamlessly integrates unsupervised Deep Image Prior (DIP) with a supervised pre-trained denoiser (DRUNet) under a unified Denoising-based Regularization (RED) optimization scheme, requiring no paired annotations. Net2Net synergistically leverages the input-specific modeling capability of untrained networks and the rich noise statistics encoded in large-scale pre-trained models, achieving strong generalization across diverse noise types and imaging conditions without compromising inference efficiency. Extensive experiments on multiple real-world denoising benchmarks demonstrate that Net2Net significantly outperforms existing state-of-the-art methods—especially under extreme data scarcity—while maintaining lightweight deployment.
This work addresses the limitations of existing deep learning approaches for RAW image denoising, which often neglect classical denoising priors, resulting in overly complex models with limited generalization. To overcome this, we propose the first learnable non-local module that explicitly embeds the classical non-local self-similarity prior into a neural network. Our method integrates multi-scale feature extraction, learnable matching and collaborative filtering, noise-level map conditioning, and joint training on both synthetic and real-world noise. The resulting model achieves performance comparable to state-of-the-art CNN- and Transformer-based methods across multiple benchmarks and real datasets, while significantly reducing parameter count and demonstrating strong cross-sensor generalization capability.
Biomedical image denoising faces bottlenecks of high computational cost and reliance on scarce clean ground-truth annotations. Method: We propose Noise2Detail, an ultra-lightweight unsupervised multi-stage denoising framework. Departing from supervised learning, it adopts a Noise2Noise-inspired self-supervised training strategy without requiring clean labels. Its core innovation lies in a noise-decoupled multi-stage pipeline: the first stage breaks spatial noise correlations to generate structural priors, while the second stage directly reconstructs fine-grained details from noisy inputs. Leveraging a highly compact network architecture and stage-wise noise separation, it achieves both real-time inference speed and high-fidelity detail preservation under minimal computational overhead. Contribution/Results: Extensive experiments demonstrate that Noise2Detail significantly outperforms existing data-free methods across diverse biomedical imaging tasks—including fluorescence microscopy, electron microscopy, and histopathology—enabling practical deployment in annotation-scarce clinical settings.
Traditional image denoising methods are constrained by specific noise types and fixed image dimensions, limiting their effectiveness in handling multi-source image degradation under complex real-world conditions. This work proposes an innovative data preprocessing strategy that abandons conventional denoising in favor of directly selecting high-quality images based on image quality assessment metrics and adaptive optimal thresholds, thereby constructing a high-fidelity training set suitable for deep learning. The approach requires no uniform image resizing and is compatible with diverse noise types and acquisition environments. Evaluated on traffic sign and general object recognition tasks, models trained with this method achieve average accuracies of 93.8% and 84.9%, respectively, significantly outperforming existing approaches and demonstrating strong efficacy and generalization capability for practical applications such as autonomous driving.
This work systematically investigates noise-based pretraining as an initialization strategy for implicit neural representations (INRs), revealing that unstructured noise—such as Gaussian or uniform distributions—significantly enhances INR generalization to unseen signals without requiring real data. Moreover, noise endowed with the characteristic 1/|f|^α spectral profile of natural images achieves a superior trade-off between signal fitting and inverse imaging tasks like image denoising. The proposed approach consistently outperforms conventional random initialization across both image and video applications, matching the performance of state-of-the-art data-driven methods while offering a practical prior for low-resource scenarios. These findings provide theoretical and empirical insights into the role of initialization in INRs, demonstrating that carefully designed synthetic noise can serve as an effective inductive bias.