🤖 AI Summary
Existing diffusion models struggle to preserve identity consistency in face restoration under severe occlusion and conflicting text guidance. To address this challenge, this work proposes ReSem-Face, a cascaded diffusion framework that introduces explicit identity semantic priors for the first time. The method distills identity features from multiple reference images and employs a multi-stream conditional architecture to synergistically fuse identity and text guidance, thereby strengthening semantic constraints during missing region reconstruction. Evaluated on CelebAHQ-IDI-5 and VGGFace2 benchmarks, ReSem-Face significantly outperforms current state-of-the-art methods, achieving high-fidelity identity preservation even under heavy occlusion while simultaneously enhancing the accuracy and consistency of text-controllable editing.
📝 Abstract
Face inpainting with diffusion models has recently achieved impressive visual quality, yet preserving identity fidelity under significant occlusion and conflicting text guidance remains a major challenge. To address this issue, we present Reference Semantic Inpainting for Face (ReSem-Face), a cascaded diffusion framework that introduces an explicit identity-conditioned semantic prior for multi-reference face inpainting. Our approach distills representative identity features from multiple references to reconstruct missing semantic regions, which then guide the diffusion process through a multi-stream conditioning architecture. This design provides strong semantic constraints when pixels are absent and stabilizes identity reconstruction while remaining compatible with prompt-driven edits. Experiments on CelebAHQ-IDI-5 and VGGFace2 demonstrate that ReSem-Face yields more reliable identity-preserving completion under severe semantic masks and improves text-controlled editing quality compared with representative baselines.