OverLay++: Dense-Overlap Layout-to-Image Generation Dataset

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bottleneck of insufficient high-density, multi-object interaction training data for layout-to-image generation in complex overlapping scenes. To this end, we construct a large-scale dataset comprising 500,000 images with an average of 6.6 objects per image, and propose a streamlined automated data generation pipeline that integrates dense detection with fine-grained semantic annotations to achieve high-quality synthesis. Compared to existing methods, our dataset increases annotation density by 1.67 times and description length sixfold, effectively bridging the data gap for complex scenes. Experimental results demonstrate that state-of-the-art models trained on this dataset exhibit significantly improved performance and accelerated convergence, validating the critical role of dense overlapping supervision in controllable image generation.
📝 Abstract
Layout-to-Image generation has made substantial progress in spatial and object-level control. However, existing methods still struggle with complex scenes containing many overlapping and interacting objects. We argue that training data is a particular bottleneck: existing datasets lack examples with dense, complex object interactions. To address this gap, we introduce OverLay++, a large-scale Layout-to-Image dataset with structurally complex scenes. OverLay++ contains approximately 500K images with an average of 6.6 objects per image, exceeding existing datasets by 1.67 times in annotation density. Beyond annotation density, OverLay++ provides rich semantic detail with object captions over six times longer than in current datasets. Our dataset generation pipeline is simple and produces dense, overlapping object annotations with rich per-object captions. Across multiple benchmarks, state-of-the-art Layout-to-Image methods trained on the OverLay++ dataset show consistent improvement and faster convergence, demonstrating the importance of dense, overlap-aware, and caption-rich supervision for controllable image generation.
Problem

Research questions and friction points this paper is trying to address.

Layout-to-Image Generation
Dense Overlap
Complex Scenes
Training Data Bottleneck
Innovation

Methods, ideas, or system contributions that make the work stand out.

Layout-to-Image Generation
Dense-Overlap Dataset
Data Generation Pipeline
Controllable Image Generation
Rich Object Captions