🤖 AI Summary
Existing diffusion posterior sampling methods often suffer from failure in high-frequency detail recovery, early-stage drift, and sensitivity to sampling schedules due to weak coupling between data consistency guidance and noise levels. To address these issues, this work proposes a noise–frequency continuum framework that constructs a sequence of intermediate posterior distributions, enforcing measurement consistency constraints within frequency bands matched to the current noise level. The approach integrates band-limited likelihood guidance with a multi-resolution consistency strategy, enabling progressive recovery of discernible high-frequency details grounded in reliable low-frequency corrections. Extensive experiments demonstrate state-of-the-art performance across tasks including super-resolution, image inpainting, and motion deblurring, with the latter achieving up to a 5 dB improvement in PSNR.
📝 Abstract
Diffusion posterior sampling solves inverse problems by combining a pretrained diffusion prior with measurement-consistency guidance, but it often fails to recover fine details because measurement terms are applied in a manner that is weakly coupled to the diffusion noise level. At high noise, data-consistency gradients computed from inaccurate estimates can be geometrically incongruent with the posterior geometry, inducing early-step drift, spurious high-frequency artifacts, plus sensitivity to schedules and ill-conditioned operators. To address these concerns, we propose a noise--frequency Continuation framework that constructs a continuous family of intermediate posteriors whose likelihood enforces measurement consistency only within a noise-dependent frequency band. This principle is instantiated with a stabilized posterior sampler that combines a diffusion predictor, band-limited likelihood guidance, and a multi-resolution consistency strategy that aggressively commits reliable coarse corrections while conservatively adopting high-frequency details only when they become identifiable. Across super-resolution, inpainting, and deblurring, our method achieves state-of-the-art performance and improves motion deblurring PSNR by up to 5 dB over strong baselines.