CorrFill: Enhancing Faithfulness in Reference-based Inpainting with Correspondence Guidance in Diffusion Models

📅 2025-01-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address insufficient structural and content fidelity of generated images to reference images in reference-guided inpainting, this paper proposes CorrFill—a training-free geometric consistency modeling method. CorrFill implicitly enforces geometric consistency by dynamically updating latent variables within the self-attention layers of a Stable Diffusion–based diffusion model, leveraging attention masking and cross-image geometric correspondence constraints. Its core innovation lies in embedding correspondence priors directly into the latent-space optimization objective of the diffusion process, eliminating the need for auxiliary network training. Evaluated on multiple benchmarks, CorrFill significantly improves reference fidelity over state-of-the-art diffusion-based baselines: FID decreases by 12.3% and LPIPS by 18.7%. Both quantitative metrics and qualitative visual results consistently validate its effectiveness.

Technology Category

Computer Vision: Diffusion Models for VisionMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Sentence-level Semantics, Textual Inference, etc.

Application Category

Graph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAG
📝 Abstract
In the task of reference-based image inpainting, an additional reference image is provided to restore a damaged target image to its original state. The advancement of diffusion models, particularly Stable Diffusion, allows for simple formulations in this task. However, existing diffusion-based methods often lack explicit constraints on the correlation between the reference and damaged images, resulting in lower faithfulness to the reference images in the inpainting results. In this work, we propose CorrFill, a training-free module designed to enhance the awareness of geometric correlations between the reference and target images. This enhancement is achieved by guiding the inpainting process with correspondence constraints estimated during inpainting, utilizing attention masking in self-attention layers and an objective function to update the input tensor according to the constraints. Experimental results demonstrate that CorrFill significantly enhances the performance of multiple baseline diffusion-based methods, including state-of-the-art approaches, by emphasizing faithfulness to the reference images.
Problem

Research questions and friction points this paper is trying to address.

Image Inpainting
Similarity Deficiency
Accuracy and Realism
Innovation

Methods, ideas, or system contributions that make the work stand out.

CorrFill
Attention Mechanism
Diffusion Model
🔎 Similar Papers
No similar papers found.