Score
Designs and implements techniques and pipelines to create composite images by blending and integrating foreground and background layers; ensures consistent lighting, color distribution, and geometric alignment across layers while producing diverse, coherent foreground appearances.
This work addresses the pervasive visual inconsistency—arising from mismatches in color, scale, shape, illumination, shadows, and reflections—between inserted objects and background scenes in image compositing. To this end, we propose the first holistic taxonomy and unified framework for image synthesis, encompassing core subtasks including object placement, scene fusion, color harmonization, and shadow/reflection generation. Methodologically, the framework integrates CNNs, GANs, diffusion models, and multi-scale feature alignment techniques. Our contributions are threefold: (1) a standardized benchmark suite unifying major datasets (e.g., iHarmony4, HCOCO); (2) libcom—an open-source, modular image compositing toolbox implementing over ten state-of-the-art algorithms; and (3) a paradigm shift toward systematic modeling and engineering-ready deployment in image synthesis. Extensive experiments demonstrate the framework’s generality, robustness, and practical utility across diverse compositional scenarios.
To address the lack of editable layered representations in professional image synthesis, this paper proposes a novel “post-generation decomposition” paradigm: instead of training layered generative models, it leverages pre-trained diffusion models to synthesize full images and then applies a generative-prior-driven unsupervised disentanglement method to decompose them intelligently into foreground and background layers. Key contributions include: (1) integrating generative prior constraints into layer decomposition to enhance semantic consistency; and (2) introducing a high-frequency feature alignment module that significantly improves edge fidelity and fine-detail accuracy. The method robustly produces high-quality layered outputs across multi-scale and multi-content scenarios, enabling downstream applications such as relighting, occlusion repair, and local editing. Crucially, it achieves an effective balance between generation quality and controllability—without requiring annotated data or additional model training.
Existing layered image synthesis methods face limitations in foreground-background separation, data availability, synthesis quality, and scene diversity. This work proposes the BFS framework, which, for the first time, transfers knowledge from non-layered image synthesis to layered generation. BFS employs a dual-branch diffusion model that jointly synthesizes a foreground layer—complete with visual effects such as shadows and reflections—and a composite image, ensuring photorealism and coherence. To address data scarcity, the method introduces a two-stage training strategy that requires only high-quality non-layered images. Experimental results and user studies demonstrate that BFS significantly outperforms current approaches in terms of synthesis quality, visual consistency, and scene diversity.
This work addresses fine-grained material editing from example images, enabling precise manipulation of object appearance and continuous control over physical material properties—including roughness, metallicness, transparency, and emissivity. Methodologically, we first identify material-attributive modules within a pre-trained diffusion UNet and introduce a novel CLIP-space material direction learning framework: a lightweight direction prediction network models semantic material directions in the CLIP embedding space, enabling module-level latent-space intervention in the UNet. This design supports multi-material blending, recombination, and cross-scene transfer in a single forward pass. Experiments demonstrate substantial improvements in material fidelity and controllability, outperforming state-of-the-art methods in both qualitative and quantitative evaluations—without requiring fine-tuning or additional training.
Existing image layer decomposition methods struggle to capture the structural and stylistic characteristics of anime illustrations, limiting controllable editing capabilities. This work proposes a structured layering framework aligned with professional anime production pipelines, decomposing illustrations into semantically meaningful layers—such as line art, flat colors, shadows, and highlights—that correspond directly to standard animation workflows. By introducing lightweight layer-specific semantic embeddings and a hierarchical supervision loss, together with the first high-quality layered dataset that emulates authentic anime production processes, our approach achieves highly accurate and visually consistent layer decomposition. The method substantially enhances flexibility in downstream editing tasks such as recoloring and texture insertion, marking the first integration of industrial-grade anime creation logic into image layer modeling.
Existing image generation methods struggle to simultaneously achieve controllability, consistency, and photorealism when editing specific elements, particularly due to inadequate modeling of physical effects such as shadows and reflections, as well as coherent composition relationships. To address this, this work proposes LASAGNA, a unified framework that jointly generates foreground objects with alpha transparency (RGBA) and photorealistic backgrounds, supporting multi-condition control via text prompts, foreground content, background context, and spatial masks. The contributions include the release of the LASAGNA-48K dataset and LASAGNABENCH—the first hierarchical editing benchmark—alongside innovations in high-fidelity RGBA compositing, physically plausible lighting and shadow modeling, and a novel training strategy. Experiments demonstrate that LASAGNA significantly outperforms existing approaches while preserving object identity and visual consistency, enabling diverse high-fidelity image editing applications.
This work addresses the challenge of preserving inter-layer structural consistency in AI-based layered image editing, where existing methods often suffer from background leakage into foreground layers and unstable alpha channels. The authors propose a training-free, context-conditioned framework for layered editing that employs a dual-stream attention mechanism to leverage contextual information from unedited layers, thereby guiding text-driven editing of the target RGBA layer while strictly preserving all other layers. This approach represents the first training-free method capable of context-aware layered editing, explicitly safeguarding layer integrity and alpha channel fidelity. To advance research in this direction, the authors also introduce LayerEditBench, a dedicated evaluation benchmark. Experiments demonstrate that the proposed method significantly outperforms strong baselines in both editing fidelity and alpha stability, effectively enhancing the realism and layer purity of composite images.
Existing diffusion-based object insertion methods treat the task solely as 2D image inpainting, lacking explicit control over the 3D pose of inserted objects. This work proposes DIRECT, a novel framework that, for the first time, integrates user-controllable 3D proxies with 2D diffusion generation. By decoupling appearance, geometry, and contextual guidance signals and injecting them into separate pathways, DIRECT effectively mitigates feature entanglement. The approach enables precise 3D pose control and scene-adaptive placement while preserving the visual fidelity of reference objects. Experimental results demonstrate that DIRECT outperforms existing methods in both geometric controllability and visual quality, supporting high-fidelity object insertion under interactive 3D pose adjustments.
本文提出一种自动管道,通过结合正负空间来优化视觉概念融合,使用语义推理和几何约束生成更富表现力和创意的图像。
Existing video editing approaches struggle to simultaneously achieve photorealistic synthesis, object reusability, and cross-frame consistency due to the absence of explicit hierarchical representations, particularly in object insertion and layer decomposition tasks where they are constrained by implicit modeling or per-scene optimization. To address this, this work proposes DBL-Diffusion, a dual-branch diffusion framework, and introduces TriLayer, the first large-scale dataset comprising RGB-foreground-background triplets. By leveraging shared denoising and cross-branch interaction mechanisms, the method jointly learns RGB videos and explicit RGBA foreground layers, enabling end-to-end learning of explicit layered video representations for the first time. Experiments demonstrate that the proposed approach significantly enhances realism and editing flexibility in object insertion while substantially improving the quality of foreground and background recovery in layer decomposition.
Existing image synthesis methods struggle to simultaneously achieve high realism—ensuring semantic and pose consistency between foreground and background—and high fidelity in preserving foreground details. This work proposes a novel two-stage generative framework that decouples these objectives for the first time: the first stage synthesizes a foreground deformation compatible with the background to guarantee realism, while the second stage refines this result to reconstruct high-fidelity details. Evaluated on the MureCOM dataset, the proposed approach significantly outperforms current single-stage methods, achieving both natural scene integration and faithful detail preservation. The code and models are publicly released.