🤖 AI Summary
To address the challenge of jointly achieving precise layout control and faithful attribute rendering in multi-instance text-to-image generation (MIG), this paper proposes a two-stage decoupled framework. In the first stage, LDM3D generates high-fidelity depth maps to enable fine-grained instance localization and holistic scene structure modeling. In the second stage, a pre-trained ControlNet performs zero-shot conditional rendering, enabling plug-and-play integration with arbitrary base diffusion models (e.g., SD2, SDXL) without fine-tuning. We introduce a novel depth-driven compositional paradigm and design a lightweight depth adapter to enhance layout controllability. Evaluated on COCO-Position and COCO-MIG benchmarks, our method achieves significant improvements in layout accuracy and attribute fidelity while demonstrating strong generalization and model-agnostic compatibility. The implementation is publicly available.
📝 Abstract
The increasing demand for controllable outputs in text-to-image generation has spurred advancements in multi-instance generation (MIG), allowing users to define both instance layouts and attributes. However, unlike image-conditional generation methods such as ControlNet, MIG techniques have not been widely adopted in state-of-the-art models like SD2 and SDXL, primarily due to the challenge of building robust renderers that simultaneously handle instance positioning and attribute rendering. In this paper, we introduce Depth-Driven Decoupled Instance Synthesis (3DIS), a novel framework that decouples the MIG process into two stages: (i) generating a coarse scene depth map for accurate instance positioning and scene composition, and (ii) rendering fine-grained attributes using pre-trained ControlNet on any foundational model, without additional training. Our 3DIS framework integrates a custom adapter into LDM3D for precise depth-based layouts and employs a finetuning-free method for enhanced instance-level attribute rendering. Extensive experiments on COCO-Position and COCO-MIG benchmarks demonstrate that 3DIS significantly outperforms existing methods in both layout precision and attribute rendering. Notably, 3DIS offers seamless compatibility with diverse foundational models, providing a robust, adaptable solution for advanced multi-instance generation. The code is available at: https://github.com/limuloo/3DIS.