Image Editing with Diffusion Models: A Survey

📅 2025-04-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
The field of diffusion model-based image editing lacks a systematic, unified taxonomy and evaluation framework. Method: This paper introduces the first four-dimensional analytical framework—encompassing task formalization, method categorization, evaluation metrics, and benchmark datasets—and innovatively classifies editing approaches into three paradigms: inversion-based, fine-tuning-based, and adapter-based methods. It further establishes a comprehensive survey of evaluation protocols and datasets covering all major editing paradigms. Contribution/Results: Based on a distilled analysis of over 100 representative works, the paper delineates performance boundaries and application scopes for each paradigm, proposes multi-granularity evaluation techniques, and unifies key technical components—including latent-space inversion, parameter-efficient fine-tuning (PEFT), and adapter architectures. The framework advances standardization, reproducibility, and systematic progress in diffusion-based image editing research.

Technology Category

Computer Vision: Diffusion Models for VisionMachine Learning: Deep Neural Architectures and Foundation ModelsGame Theory and Economic Paradigms: Adversarial Learning

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
With deeper exploration of diffusion model, developments in the field of image generation have triggered a boom in image creation. As the quality of base-model generated images continues to improve, so does the demand for further application like image editing. In recent years, many remarkable works are realizing a wide variety of editing effects. However, the wide variety of editing types and diverse editing approaches have made it difficult for researchers to establish a comprehensive view of the development of this field. In this survey, we summarize the image editing field from four aspects: tasks definition, methods classification, results evaluation and editing datasets. First, we provide a definition of image editing, which in turn leads to a variety of editing task forms from the perspective of operation parts and manipulation actions. Subsequently, we categorize and summary methods for implementing editing into three categories: inversion-based, fine-tuning-based and adapter-based. In addition, we organize the currently used metrics, available datasets and corresponding construction methods. At the end, we present some visions for the future development of the image editing field based on the previous summaries.
Problem

Research questions and friction points this paper is trying to address.

Summarize diverse image editing tasks and methods
Classify editing approaches into three main categories
Organize evaluation metrics and datasets for editing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Surveying image editing with diffusion models
Classifying methods into three categories
Organizing metrics and datasets comprehensively
🔎 Similar Papers
2024-09-16Philosophical transactions. Series A, Mathematical, physical, and engineering sciencesCitations: 8
J
Jia Wang
University of Chinese Academy of Sciences, Beijing, China; Meituan, Beijing, China
J
Jie Hu
Key Laboratory of System Software (Chinese Academy of Sciences) and State Key Laboratory of Computer Science, Institute of Software, Chinese Academy of Sciences, Beijing, China; Meituan, Beijing, China
Xiaoqi Ma
Xiaoqi Ma
Meituan, Beijing, China
H
Hanghang Ma
Meituan, Beijing, China
Xiaoming Wei
Xiaoming Wei
Meituan
computer visionmachine learning
E
Enhua Wu
Key Laboratory of System Software (Chinese Academy of Sciences) and State Key Laboratory of Computer Science, Institute of Software, Chinese Academy of Sciences, Beijing, China