AttenCraft: Attention-guided Disentanglement of Multiple Concepts for Text-to-Image Customization

📅 2024-05-28
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of multi-concept disentanglement, asynchronous feature fusion and learning, and poor structural alignment in single-image, text-to-image customization—particularly without manual masks. We propose a diffusion-based framework that enables mask-free, concept-aware image generation. Our method introduces: (1) an attention-driven, single-step automatic masking mechanism that leverages self- and cross-attention maps for concept-level segmentation; (2) Uniform and Reweighted sampling strategies to mitigate temporal inconsistency during multi-concept feature extraction; and (3) native support for complex scenarios involving three or more concepts. Extensive experiments demonstrate that our approach preserves strong text-image alignment while significantly improving structural fidelity and concept disentanglement on single images. Ablation studies confirm the effectiveness and robustness of each component, establishing state-of-the-art performance in mask-free, multi-concept customization tasks.

Technology Category

Computer Vision: Diffusion Models for VisionNatural Language Processing: GenerationMachine Learning: Deep Generative Models & Autoencoders

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGUser Modeling, Personalization and Recommendation: Federated recommendation systems and personalization
📝 Abstract
With the unprecedented performance being achieved by text-to-image (T2I) diffusion models, T2I customization further empowers users to tailor the diffusion model to new concepts absent in the pre-training dataset, termed subject-driven generation. Moreover, extracting several new concepts from a single image enables the model to learn multiple concepts, and simultaneously decreases the difficulties of training data preparation, urging the disentanglement of multiple concepts to be a new challenge. However, existing models for disentanglement commonly require pre-determined masks or retain background elements. To this end, we propose an attention-guided method, AttenCraft, for multiple concept disentanglement. In particular, our method leverages self-attention and cross-attention maps to create accurate masks for each concept within a single initialization step, omitting any required mask preparation by humans or other models. The created masks are then applied to guide the cross-attention activation of each target concept during training and achieve concept disentanglement. Additionally, we introduce Uniform sampling and Reweighted sampling schemes to alleviate the non-synchronicity of feature acquisition from different concepts, and improve generation quality. Our method outperforms baseline models in terms of image-alignment, and behaves comparably on text-alignment. Finally, we showcase the applicability of AttenCraft to more complicated settings, such as an input image containing three concepts. The project is available at https://github.com/junjie-shentu/AttenCraft.
Problem

Research questions and friction points this paper is trying to address.

Disentangles multiple concepts from single images
Addresses feature fusion and asynchronous learning issues
Enables balanced concept acquisition without manual masks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Attention maps generate concept masks automatically
Adaptive algorithm balances concept learning ratios
Feature-retaining framework prevents concept fusion
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Durham University
J
Junjie Shentu
Department of Computer Science, Durham University, UK
M
Matthew Watson
Department of Computer Science, Durham University, UK
N
N. A. Moubayed
Department of Computer Science, Durham University, UK