🤖 AI Summary
This study addresses the challenge of insufficient reference attention allocation in diffusion-based visual editing, which hinders faithful reproduction of reference images. To overcome this limitation, we propose RefGAP, a training-free correction framework that dynamically evaluates reference attention quality during the forward pass. By employing two global coefficients to adaptively adjust logit offsets, RefGAP enhances reference feature utilization without requiring method-specific intensity sweeps for hyperparameter tuning. Extensive experiments demonstrate that the proposed approach significantly improves identity fidelity in head and face swapping across seven image and video editing models. Furthermore, its strong generalization capability is validated on diverse tasks such as virtual try-on. Overall, RefGAP provides an efficient, versatile, and plug-and-play solution for reference-guided editing in diffusion models.
📝 Abstract
Reference-guided diffusion editors struggle to faithfully reproduce user-provided references. We identify a potential bottleneck in diffusion editors: many methods provide limited reference-attention allocation. For example, in LoomVideo, edit-region queries assign less than 1% of their attention mass to the reference. We introduce RefGAP, a training-free correction that determines logit-offset magnitudes online at each layer from the reference-attention mass measured during the forward pass. Positive offsets to reference logits strengthen reference usage by edit-region queries, while negative offsets for keep-region queries limit reference-induced changes outside the edit. Two global coefficients control the correction; they are selected once on validation data from four development diffusion editors and held fixed. Across seven diffusion-based image/video editors, RefGAP improves identity fidelity in head swapping and face swapping. RefGAP achieves a fidelity-preservation trade-off comparable to separately tuned constant edit-side biases, without per-approach strength sweeps. Additional experiments on virtual try-on and background replacement evaluate transfer beyond identity editing.