🤖 AI Summary
Existing learning-based image denoising methods rely on fixed noise priors, leading to poor generalization under varying real-world noise distributions. To address this, we propose a novel paradigm that decouples noise priors from image priors, and introduce a conditional optimization framework capable of estimating sensor-level noise priors directly from a single sRGB noisy image. Our key contributions are: (1) the first explicit modeling and incorporation of noise priors into denoising architectures; (2) a lightweight Local Noise Prior Estimator (LoNPE) network for pixel-wise noise prior estimation; and (3) a Conditional Denoising Transformer (CondFormer) that dynamically injects estimated noise priors into the denoising subspace via conditional self-attention. Extensive experiments demonstrate significant improvements over state-of-the-art methods on both synthetic and real-world datasets, with strong cross-device robustness and generalization capability. The code is publicly available.
📝 Abstract
Existing learning-based denoising methods typically train models to generalize the image prior from large-scale datasets, suffering from the variability in noise distributions encountered in real-world scenarios. In this work, we propose a new perspective on the denoising challenge by highlighting the distinct separation between noise and image priors. This insight forms the basis for our development of conditional optimization framework, designed to overcome the constraints of traditional denoising framework. To this end, we introduce a Locally Noise Prior Estimation (LoNPE) algorithm, which accurately estimates the noise prior directly from a single raw noisy image. This estimation acts as an explicit prior representation of the camera sensor's imaging environment, distinct from the image prior of scenes. Additionally, we design an auxiliary learnable LoNPE network tailored for practical application to sRGB noisy images. Leveraging the estimated noise prior, we present a novel Conditional Denoising Transformer (Condformer), by incorporating the noise prior into a conditional self-attention mechanism. This integration allows the Condformer to segment the optimization process into multiple explicit subspaces, significantly enhancing the model's generalization and flexibility. Extensive experimental evaluations on both synthetic and real-world datasets, demonstrate that the proposed method achieves superior performance over current state-of-the-art methods. The source code is available at https://github.com/YuanfeiHuang/Condformer.