3DIS: Depth-Driven Decoupled Instance Synthesis for Text-to-Image Generation

📅 2024-10-16
🏛️ arXiv.org
📈 Citations: 21
✨ Influential: 1
📄 PDF
🤖 AI Summary
To address the challenge of jointly achieving precise layout control and faithful attribute rendering in multi-instance text-to-image generation (MIG), this paper proposes a two-stage decoupled framework. In the first stage, LDM3D generates high-fidelity depth maps to enable fine-grained instance localization and holistic scene structure modeling. In the second stage, a pre-trained ControlNet performs zero-shot conditional rendering, enabling plug-and-play integration with arbitrary base diffusion models (e.g., SD2, SDXL) without fine-tuning. We introduce a novel depth-driven compositional paradigm and design a lightweight depth adapter to enhance layout controllability. Evaluated on COCO-Position and COCO-MIG benchmarks, our method achieves significant improvements in layout accuracy and attribute fidelity while demonstrating strong generalization and model-agnostic compatibility. The implementation is publicly available.

Technology Category

Computer Vision: Diffusion Models for VisionMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: LLM based quality controls for crowd work
📝 Abstract
The increasing demand for controllable outputs in text-to-image generation has spurred advancements in multi-instance generation (MIG), allowing users to define both instance layouts and attributes. However, unlike image-conditional generation methods such as ControlNet, MIG techniques have not been widely adopted in state-of-the-art models like SD2 and SDXL, primarily due to the challenge of building robust renderers that simultaneously handle instance positioning and attribute rendering. In this paper, we introduce Depth-Driven Decoupled Instance Synthesis (3DIS), a novel framework that decouples the MIG process into two stages: (i) generating a coarse scene depth map for accurate instance positioning and scene composition, and (ii) rendering fine-grained attributes using pre-trained ControlNet on any foundational model, without additional training. Our 3DIS framework integrates a custom adapter into LDM3D for precise depth-based layouts and employs a finetuning-free method for enhanced instance-level attribute rendering. Extensive experiments on COCO-Position and COCO-MIG benchmarks demonstrate that 3DIS significantly outperforms existing methods in both layout precision and attribute rendering. Notably, 3DIS offers seamless compatibility with diverse foundational models, providing a robust, adaptable solution for advanced multi-instance generation. The code is available at: https://github.com/limuloo/3DIS.
Problem

Research questions and friction points this paper is trying to address.

Generates multi-instance images from text with precise layouts
Decouples layout and attribute rendering to improve control
Enables compatibility with existing models without extra training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decouples multi-instance generation into two stages
Uses depth map for layout and ControlNet for attributes
Compatible with various models without extra training
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.