Score
Designs and builds composite raster images by combining multiple photographic or rendered elements using Photoshop tools—layers, masks, selections, blending modes, retouching, color grading, and filters—to create seamless, realistic, or stylized scenes and visual effects. Implements non‑destructive workflows and prepares images for intended outputs by managing resolution, color spaces, and export settings.
This work addresses the pervasive visual inconsistency—arising from mismatches in color, scale, shape, illumination, shadows, and reflections—between inserted objects and background scenes in image compositing. To this end, we propose the first holistic taxonomy and unified framework for image synthesis, encompassing core subtasks including object placement, scene fusion, color harmonization, and shadow/reflection generation. Methodologically, the framework integrates CNNs, GANs, diffusion models, and multi-scale feature alignment techniques. Our contributions are threefold: (1) a standardized benchmark suite unifying major datasets (e.g., iHarmony4, HCOCO); (2) libcom—an open-source, modular image compositing toolbox implementing over ten state-of-the-art algorithms; and (3) a paradigm shift toward systematic modeling and engineering-ready deployment in image synthesis. Extensive experiments demonstrate the framework’s generality, robustness, and practical utility across diverse compositional scenarios.
To address the irreversibility of layer editing in rasterized image generation, this paper proposes an invertible layer decomposition method. It iteratively extracts occlusion-free foreground layers and introduces a novel quality metric grounded in layer-wise visual consistency—commonly observed in professional graphic design—to alleviate the ill-posedness of decomposition. The method adopts a two-stage strategy: “progressive extraction” followed by “refinement optimization,” integrating domain-specific priors with state-of-the-art generative models (e.g., diffusion models) for layer reconstruction and validation. Experiments demonstrate substantial improvements over existing baselines across multiple benchmarks, yielding layers with high fidelity and precise editability. Furthermore, the method has been integrated into mainstream image generation and layer-editing toolchains, enabling robust, re-editable creative workflows.
This work addresses the limitation of existing generative models in producing raster images without editable layer structures, which hinders downstream graphic editing. The authors propose a hybrid generation framework that, for the first time, parses text regions into re-renderable protocols and integrates a vision-language model with an RGBA multi-branch diffusion architecture to separately reconstruct text, background, and sticker layers. To better align with human design preferences, they introduce ParserReward and Group Relative Policy Optimization—two reinforcement learning mechanisms—that substantially enhance controllability and editing flexibility. Evaluated on the Parser-40K and Crello datasets, the method achieves an average performance gain of 23.7% over current state-of-the-art approaches.
Existing image layer decomposition methods struggle to capture the structural and stylistic characteristics of anime illustrations, limiting controllable editing capabilities. This work proposes a structured layering framework aligned with professional anime production pipelines, decomposing illustrations into semantically meaningful layers—such as line art, flat colors, shadows, and highlights—that correspond directly to standard animation workflows. By introducing lightweight layer-specific semantic embeddings and a hierarchical supervision loss, together with the first high-quality layered dataset that emulates authentic anime production processes, our approach achieves highly accurate and visually consistent layer decomposition. The method substantially enhances flexibility in downstream editing tasks such as recoloring and texture insertion, marking the first integration of industrial-grade anime creation logic into image layer modeling.
Current raster image synthesis is irreversible, hindering layer-level editing; existing matting and inpainting methods suffer from limited segmentation accuracy and controllability. To address these challenges, we propose LayerDecompose-DiT—a novel framework integrating Multi-Layer Conditional Adapters with a diffusion-Transformer-based multi-layer token control mechanism, enabling fine-grained, editable RGBA layer decomposition and reconstruction. We introduce the first design-oriented multi-layer decomposition benchmark dataset and corresponding evaluation metrics. Experiments demonstrate that our method significantly outperforms state-of-the-art approaches in decomposition accuracy, semantic consistency, and editing controllability. Critically, the generated layers are directly importable into mainstream design tools (e.g., PowerPoint), bridging theoretical innovation with practical applicability.
This paper addresses three key challenges in raster-to-vector conversion: weak semantic alignment, hierarchical redundancy, and low visual fidelity. To this end, we propose a progressive hierarchical vectorization framework that first constructs a semantically aligned macro-structure and then progressively refines details across layers, yielding compact, layered vector representations. Our core contributions are: (1) a semantic simplification mechanism leveraging the feature-averaging effect of Score Distillation Sampling for progressive structural simplification; (2) a two-stage “structure-first, detail-later” vectorization paradigm; and (3) an inter-layer optimization strategy enforcing dual alignment—explicit structural alignment and implicit texture/style alignment. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art approaches across diverse image categories, producing SVGs with fewer layers, stronger semantic consistency, higher visual fidelity, and superior editability.
Existing image generation methods produce flat outputs that are difficult to edit, while real-world layered data is scarce and non-scalable, hindering the practical deployment of layer decomposition techniques. This work proposes a data-driven approach based on SynLayers, a purely synthetic dataset, operating within the CLD framework by integrating vision-language models (VLMs) to generate textual supervision alongside bounding box inputs. It demonstrates for the first time that purely synthetic data can effectively substitute real data for layered design decomposition. The method overcomes data scalability limitations, enables balanced control over the distribution of layer counts, and outperforms non-scalable alternatives such as PrismLayersPro. Experiments show that model performance saturates at around 50,000 samples, significantly alleviating the layer count imbalance problem.
This work addresses the challenge of preserving inter-layer structural consistency in AI-based layered image editing, where existing methods often suffer from background leakage into foreground layers and unstable alpha channels. The authors propose a training-free, context-conditioned framework for layered editing that employs a dual-stream attention mechanism to leverage contextual information from unedited layers, thereby guiding text-driven editing of the target RGBA layer while strictly preserving all other layers. This approach represents the first training-free method capable of context-aware layered editing, explicitly safeguarding layer integrity and alpha channel fidelity. To advance research in this direction, the authors also introduce LayerEditBench, a dedicated evaluation benchmark. Experiments demonstrate that the proposed method significantly outperforms strong baselines in both editing fidelity and alpha stability, effectively enhancing the realism and layer purity of composite images.
Existing image generation methods struggle to achieve fine-grained control over shape, color, and semantics at the element level. This work proposes a hierarchical, controllable image generation framework based on simplified vector graphics (VG). The approach first parses an image into a semantically aligned and structurally coherent hierarchical VG representation, then employs a VG-guided noise prediction mechanism to seamlessly map user edits on vector elements to high-fidelity images. By leveraging simplified VG as a novel guidance signal for photorealistic image synthesis, this method enables unprecedented fine-grained controllability across geometric, chromatic, and semantic dimensions, establishing a new paradigm for editable generative modeling. Experiments demonstrate that the framework offers intuitive, efficient, and high-quality performance in object-level manipulation, local editing, and content creation tasks.
Existing image simplification methods often rely on non-photorealistic rendering, struggling to balance visual abstraction with photorealistic fidelity. This work proposes a progressive semantic image simplification framework that iteratively reduces scene complexity through a sequence of selection, removal, and verification steps while preserving photometric realism. For the first time, it enables controllable, semantics-driven progressive simplification: a vision-language model identifies and ranks content elements by importance; generative editing combined with a learned validator ensures perceptual realism; and knowledge distillation yields an end-to-end image-to-video simplification model. The resulting simplification sequences are visually coherent and naturally support applications such as content-aware decluttering, semantic layering, and interactive editing.
Existing layered image synthesis methods face limitations in foreground-background separation, data availability, synthesis quality, and scene diversity. This work proposes the BFS framework, which, for the first time, transfers knowledge from non-layered image synthesis to layered generation. BFS employs a dual-branch diffusion model that jointly synthesizes a foreground layer—complete with visual effects such as shadows and reflections—and a composite image, ensuring photorealism and coherence. To address data scarcity, the method introduces a two-stage training strategy that requires only high-quality non-layered images. Experimental results and user studies demonstrate that BFS significantly outperforms current approaches in terms of synthesis quality, visual consistency, and scene diversity.