🤖 AI Summary
To address the limited performance of real-image denoising under extremely low-light conditions, this work introduces, for the first time, natural-language scene descriptions provided by photographers as explicit semantic priors into the denoising pipeline, moving beyond conventional purely data-driven paradigms. Methodologically, we propose a text-guided diffusion model that jointly leverages a CLIP text encoder and a U-Net-based denoising network to enable cross-modal conditional reconstruction. Evaluated on both synthetic and real low-light datasets, our approach achieves significant improvements in PSNR and SSIM; notably, in single-frame scenarios with extreme noise, structural and textural details are visibly restored. This work pioneers a language-prior-driven paradigm for real-image denoising, establishing an interpretable and controllable semantic guidance framework for low-light visual reconstruction.
📝 Abstract
Image reconstruction from noisy sensor measurements is challenging and many methods have been proposed for it. Yet, most approaches focus on learning robust natural image priors while modeling the scene's noise statistics. In extremely low-light conditions, these methods often remain insufficient. Additional information is needed, such as multiple captures or, as suggested here, scene description. As an alternative, we propose using a text-based description of the scene as an additional prior, something the photographer can easily provide. Inspired by the remarkable success of text-guided diffusion models in image generation, we show that adding image caption information significantly improves image denoising and reconstruction for both synthetic and real-world images.