๐ค AI Summary
To address the challenges of synthesizing complex scenes and achieving fine-grained style control in layout-to-image (L2I) generation, this paper proposes a controllable diffusion-based generation framework. Our method introduces two key innovations: (1) Edge-Aware Normalization (EA-Norm), which incorporates edge-structure priors into the denoising process to enhance layout fidelity; and (2) Styled-Mask Attention, which explicitly models object-level semantic relationships while enabling independent, disentangled style controlโensuring both inter-object consistency and per-object stylistic customization. To our knowledge, this is the first diffusion-based L2I approach that jointly models global layout constraints and local semantic-style interactions. Extensive experiments on COCO-Stuff and Visual Genome benchmarks demonstrate significant improvements over state-of-the-art methods across layout fidelity, style controllability, and relational reasoning capability. Generated images exhibit higher visual fidelity and enhanced diversity.
๐ Abstract
In layout-to-image (L2I) synthesis, controlled complex scenes are generated from coarse information like bounding boxes. Such a task is exciting to many downstream applications because the input layouts offer strong guidance to the generation process while remaining easily reconfigurable by humans. In this paper, we proposed STyled LAYout Diffusion (STAY Diffusion), a diffusion-based model that produces photo-realistic images and provides fine-grained control of stylized objects in scenes. Our approach learns a global condition for each layout, and a self-supervised semantic map for weight modulation using a novel Edge-Aware Normalization (EA Norm). A new Styled-Mask Attention (SM Attention) is also introduced to cross-condition the global condition and image feature for capturing the objects' relationships. These measures provide consistent guidance through the model, enabling more accurate and controllable image generation. Extensive benchmarking demonstrates that our STAY Diffusion presents high-quality images while surpassing previous state-of-the-art methods in generation diversity, accuracy, and controllability.