🤖 AI Summary
To address insufficient structural and content fidelity of generated images to reference images in reference-guided inpainting, this paper proposes CorrFill—a training-free geometric consistency modeling method. CorrFill implicitly enforces geometric consistency by dynamically updating latent variables within the self-attention layers of a Stable Diffusion–based diffusion model, leveraging attention masking and cross-image geometric correspondence constraints. Its core innovation lies in embedding correspondence priors directly into the latent-space optimization objective of the diffusion process, eliminating the need for auxiliary network training. Evaluated on multiple benchmarks, CorrFill significantly improves reference fidelity over state-of-the-art diffusion-based baselines: FID decreases by 12.3% and LPIPS by 18.7%. Both quantitative metrics and qualitative visual results consistently validate its effectiveness.
📝 Abstract
In the task of reference-based image inpainting, an additional reference image is provided to restore a damaged target image to its original state. The advancement of diffusion models, particularly Stable Diffusion, allows for simple formulations in this task. However, existing diffusion-based methods often lack explicit constraints on the correlation between the reference and damaged images, resulting in lower faithfulness to the reference images in the inpainting results. In this work, we propose CorrFill, a training-free module designed to enhance the awareness of geometric correlations between the reference and target images. This enhancement is achieved by guiding the inpainting process with correspondence constraints estimated during inpainting, utilizing attention masking in self-attention layers and an objective function to update the input tensor according to the constraints. Experimental results demonstrate that CorrFill significantly enhances the performance of multiple baseline diffusion-based methods, including state-of-the-art approaches, by emphasizing faithfulness to the reference images.