raster image editing

Creates and modifies pixel-based (raster) images using tools and techniques—commonly in Adobe Photoshop—such as layers, masks, selections, retouching, compositing, color correction, and digital painting, and prepares and exports optimized raster assets at appropriate resolution and color modes for print or screen.

rasterimageediting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.83
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$164K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

LayerD: Decomposing Raster Graphic Designs into Layers

Sep 29, 2025
TS
Tomoyuki Suzuki
🏛️ CyberAgent | Tohoku University

To address the irreversibility of layer editing in rasterized image generation, this paper proposes an invertible layer decomposition method. It iteratively extracts occlusion-free foreground layers and introduces a novel quality metric grounded in layer-wise visual consistency—commonly observed in professional graphic design—to alleviate the ill-posedness of decomposition. The method adopts a two-stage strategy: “progressive extraction” followed by “refinement optimization,” integrating domain-specific priors with state-of-the-art generative models (e.g., diffusion models) for layer reconstruction and validation. Experiments demonstrate substantial improvements over existing baselines across multiple benchmarks, yielding layers with high fidelity and precise editability. Furthermore, the method has been integrated into mainstream image generation and layer-editing toolchains, enabling robust, re-editable creative workflows.

Decomposing raster graphics into editable layersEnabling re-editable workflows from composite imagesExtracting uniform appearance layers iteratively

This work addresses the limitation of existing generative models in producing raster images without editable layer structures, which hinders downstream graphic editing. The authors propose a hybrid generation framework that, for the first time, parses text regions into re-renderable protocols and integrates a vision-language model with an RGBA multi-branch diffusion architecture to separately reconstruct text, background, and sticker layers. To better align with human design preferences, they introduce ParserReward and Group Relative Policy Optimization—two reinforcement learning mechanisms—that substantially enhance controllability and editing flexibility. Evaluated on the Parser-40K and Crello datasets, the method achieves an average performance gain of 23.7% over current state-of-the-art approaches.

diffusion architectureeditable layersgraphic design parsing

Controllable Layer Decomposition for Reversible Multi-Layer Image Generation

Nov 20, 2025
ZL
Zihao Liu
🏛️ Tsinghua University | Sun Yat-sen University

Current raster image synthesis is irreversible, hindering layer-level editing; existing matting and inpainting methods suffer from limited segmentation accuracy and controllability. To address these challenges, we propose LayerDecompose-DiT—a novel framework integrating Multi-Layer Conditional Adapters with a diffusion-Transformer-based multi-layer token control mechanism, enabling fine-grained, editable RGBA layer decomposition and reconstruction. We introduce the first design-oriented multi-layer decomposition benchmark dataset and corresponding evaluation metrics. Experiments demonstrate that our method significantly outperforms state-of-the-art approaches in decomposition accuracy, semantic consistency, and editing controllability. Critically, the generated layers are directly importable into mainstream design tools (e.g., PowerPoint), bridging theoretical innovation with practical applicability.

Achieving reversible multi-layer image separation for raster imagesEnabling direct layer manipulation in standard design toolsOvercoming limitations in layer decomposition controllability and precision

Layered Image Vectorization via Semantic Simplification

Jun 08, 2024
ZW
Zhenyu Wang
🏛️ Shenzhen University | Tel-Aviv University

This paper addresses three key challenges in raster-to-vector conversion: weak semantic alignment, hierarchical redundancy, and low visual fidelity. To this end, we propose a progressive hierarchical vectorization framework that first constructs a semantically aligned macro-structure and then progressively refines details across layers, yielding compact, layered vector representations. Our core contributions are: (1) a semantic simplification mechanism leveraging the feature-averaging effect of Score Distillation Sampling for progressive structural simplification; (2) a two-stage “structure-first, detail-later” vectorization paradigm; and (3) an inter-layer optimization strategy enforcing dual alignment—explicit structural alignment and implicit texture/style alignment. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art approaches across diverse image categories, producing SVGs with fewer layers, stronger semantic consistency, higher visual fidelity, and superior editability.

Achieves high semantic alignment and compact layered representation.Progressive image vectorization from semantic-aligned structures to details.Two-stage vectorization: structural buildup and visual refinement.

PixPerfect: Seamless Latent Diffusion Local Editing with Discriminative Pixel-Space Refinement

Dec 02, 2025
HZ
Haitian Zheng
🏛️ Adobe Research | University of Rochester

Latent diffusion models (LDMs) suffer from pixel-level inconsistencies—including color shifts, texture mismatches, and boundary seams—in local image editing due to aggressive latent-space compression. Existing approaches struggle to balance generality and fidelity. To address this, we propose a universal, pixel-level refinement framework comprising three key components: (1) a differentiable discriminative pixel-space modeling module that explicitly enforces fine-grained detail consistency; (2) a background-aware latent decoding mechanism coupled with adversarial artifact simulation during training to suppress structural artifacts; and (3) a plug-and-play direct pixel-space optimization module compatible with diverse implicit representations and editing tasks. Extensive experiments on image inpainting, object removal, and object insertion demonstrate substantial improvements in visual fidelity. Our method achieves state-of-the-art performance across multiple benchmarks while exhibiting strong generalization and practical applicability.

Addresses pixel-level inconsistencies in latent diffusion image editingEliminates chromatic shifts, texture mismatches, and visible seamsEnsures seamless local edits across diverse LDM architectures and tasks

Latest Papers

What's happening recently
View more

Current image generation models produce raster images that are difficult to edit, as their content is flattened into pixels, precluding inspection, modification, or reuse of individual components. This work introduces the novel task of “image-to-editable reconstruction” and proposes DrawAI-Flow, a two-stage agent-based workflow that recovers structured, directly manipulable graphic representations from raster images while preserving visual fidelity. We establish DrawAI-Bench, the first benchmark for this task, and design an iterative pipeline integrating parsing and reconstruction agents. Our approach employs a multimodal agent architecture coupled with a code rendering–verification–correction loop to generate executable graphic code. Evaluation across 13 models and five agent frameworks on DrawAI-Bench demonstrates that DrawAI-Flow significantly improves reconstruction quality, highlighting the critical roles of model capabilities, framework selection, and workflow design.

editable artifact generationimage-to-editable reconstructionraster image editing

This work addresses the challenge of preserving inter-layer structural consistency in AI-based layered image editing, where existing methods often suffer from background leakage into foreground layers and unstable alpha channels. The authors propose a training-free, context-conditioned framework for layered editing that employs a dual-stream attention mechanism to leverage contextual information from unedited layers, thereby guiding text-driven editing of the target RGBA layer while strictly preserving all other layers. This approach represents the first training-free method capable of context-aware layered editing, explicitly safeguarding layer integrity and alpha channel fidelity. To advance research in this direction, the authors also introduce LayerEditBench, a dedicated evaluation benchmark. Experiments demonstrate that the proposed method significantly outperforms strong baselines in both editing fidelity and alpha stability, effectively enhancing the realism and layer purity of composite images.

alpha channel stabilitycross-layer consistencylayered image editing

This work addresses the inefficiencies in 2D graphics rasterization caused by the semantic complexity and opaque execution model of existing libraries such as Skia, which often lead applications to emit suboptimal drawing commands. We present μSkia, the first formal semantics for Skia, mechanized in Lean and structured in three layers to precisely model canvas state, layer stacks, blending, and color filters, thereby achieving separation of concerns and extensibility. Building upon this semantics, we develop a high-performance optimizer that guarantees semantic correctness of transformations through pattern matching and translation validation, while uncovering several implicit side-effect conditions. Evaluated on 99 real-world Skia programs extracted from the Top 100 websites, the optimizer achieves an average speedup of 18.7% on modern GPU backends, with each optimization taking less than 32 microseconds and demonstrating consistent performance gains across multiple platforms.

execution modelgraphics libraryperformance inefficiency

Existing image simplification methods often rely on non-photorealistic rendering, struggling to balance visual abstraction with photorealistic fidelity. This work proposes a progressive semantic image simplification framework that iteratively reduces scene complexity through a sequence of selection, removal, and verification steps while preserving photometric realism. For the first time, it enables controllable, semantics-driven progressive simplification: a vision-language model identifies and ranks content elements by importance; generative editing combined with a learned validator ensures perceptual realism; and knowledge distillation yields an end-to-end image-to-video simplification model. The resulting simplification sequences are visually coherent and naturally support applications such as content-aware decluttering, semantic layering, and interactive editing.

Content RemovalImage SimplificationPhotorealistic Simplification

Hot Scholars

XC

Xuhang Chen

Huizhou University
computational imaginglow-level visioncomputational photography