🤖 AI Summary
The field of diffusion model-based image editing lacks a systematic, unified taxonomy and evaluation framework. Method: This paper introduces the first four-dimensional analytical framework—encompassing task formalization, method categorization, evaluation metrics, and benchmark datasets—and innovatively classifies editing approaches into three paradigms: inversion-based, fine-tuning-based, and adapter-based methods. It further establishes a comprehensive survey of evaluation protocols and datasets covering all major editing paradigms. Contribution/Results: Based on a distilled analysis of over 100 representative works, the paper delineates performance boundaries and application scopes for each paradigm, proposes multi-granularity evaluation techniques, and unifies key technical components—including latent-space inversion, parameter-efficient fine-tuning (PEFT), and adapter architectures. The framework advances standardization, reproducibility, and systematic progress in diffusion-based image editing research.
📝 Abstract
With deeper exploration of diffusion model, developments in the field of image generation have triggered a boom in image creation. As the quality of base-model generated images continues to improve, so does the demand for further application like image editing. In recent years, many remarkable works are realizing a wide variety of editing effects. However, the wide variety of editing types and diverse editing approaches have made it difficult for researchers to establish a comprehensive view of the development of this field. In this survey, we summarize the image editing field from four aspects: tasks definition, methods classification, results evaluation and editing datasets. First, we provide a definition of image editing, which in turn leads to a variety of editing task forms from the perspective of operation parts and manipulation actions. Subsequently, we categorize and summary methods for implementing editing into three categories: inversion-based, fine-tuning-based and adapter-based. In addition, we organize the currently used metrics, available datasets and corresponding construction methods. At the end, we present some visions for the future development of the image editing field based on the previous summaries.