🤖 AI Summary
This study addresses the semantic misalignment among text, foreground, and background in personalized diffusion model generation caused by frequency band entanglement. To this end, we propose the Dual-FDM framework, which innovatively introduces a frequency-aware mechanism and a frequency-domain masking strategy to decouple dual reference images. By selectively replacing mid- and low-frequency bands, our method effectively disentangles foreground customization from background style interference. Furthermore, it integrates cross-attention reconstruction to enable simultaneous optimization of customized content and style transfer. Experimental results demonstrate that Dual-FDM significantly outperforms existing state-of-the-art methods in both generation fidelity and overall image quality.
📝 Abstract
Personalized image generation aims to synthesize text-driven images conditioned on reference images, while mainly casting the generation as image customization for foreground and style transfer for background. Previous arts of diffusion models suffers from the text misalignment with background for image customization and foreground for style transfer during the denoising process. Such facts, as we observed, rooted from the entanglement among hybrid frequency bands during the denoising process. To address such salient limitation, in this paper, we study personalized generation based on dual references - customization and color and style reference - and propose a paradigm to disentangle these Dual image references within Frequency-aware Diffusion Models, dubbed Dual-FDM, to simultaneously tackle two crucial personalized image generation tasks: customization style transfer and color style transfer, by disentangling different frequency bands via mask strategy within frequency domain. For customization style transfer, we replace the mid-frequency band of the background in the style reference with that from the foreground of the customized reference. For color style transfer, we substitute the low-frequency band of the background in the style reference with that from both the foreground and background of the color reference. Both the substituted frequency bands are used as the key and value to reconstruct the query foreground and background of the denoised personalized image.Extensive experiments validate the superiority of Dual-FDM over the state-of-the-art diffusion models for personalized image generation. Our code can be accessed from https://github.com/htyjers/Dual-FDM.