Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications

📅 2026-07-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses instruction-based image editing (IIE)—enabling precise and controllable image manipulation through natural language commands. To advance this emerging field, we establish a unified conceptual framework encompassing task formulation, data curation, model architectures, evaluation protocols, and real-world applicability. We further introduce CDD-IIE Bench, the first comprehensive benchmark for IIE, which facilitates multi-dimensional and fine-grained performance diagnosis. By systematically integrating techniques from GANs, diffusion models, autoregressive models, and large language/vision-language models, we conduct an extensive empirical comparison of prominent open-source methods, elucidating their respective strengths and limitations. Our analysis clarifies key evolutionary trajectories in IIE methodologies and provides the community with a standardized evaluation toolkit and actionable directions for future research.
📝 Abstract
Instruction-based Image Editing (IIE) aims to transform a given image into a new one based on textual instructions. Advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) have accelerated progress toward practical ``one-sentence image editing" systems. This survey presents a systematic taxonomy and comprehensive review of IIE research, structured around five core dimensions: (1) task definition and hierarchical categorization of editing operations, (2) methodologies for training data construction, (3) architectural evolution from GAN-based to diffusion and autoregressive paradigms, (4) standardized evaluation metrics and benchmark development, and (5) introduction of commercial solutions. Our analysis shows critical technological milestones across model generations. We further propose a Comprehensive, in-Depth, and Diagnostic benchmark for IIE task (CDD-IIE Bench), which can rigorously assess the multiple aspects of model performance. Through empirical comparisons of open-source solutions, we highlight their respective capabilities and limitations. Finally, we discuss future research directions to advance the field.
Problem

Research questions and friction points this paper is trying to address.

Instruction-based Image Editing
Text-to-Image Editing
Vision-Language Models
Image Manipulation
Multimodal Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Instruction-based Image Editing
CDD-IIE Bench
Vision-Language Models
Diffusion Models
Benchmark Evaluation
🔎 Similar Papers
No similar papers found.
X
Xianghao Zang
Institute of Artificial Intelligence (TeleAI), China Telecom, Beijing, China
Z
Zijian Jiang
Institute of Artificial Intelligence (TeleAI), China Telecom, Beijing, China
J
Jiarong Cheng
Institute of Artificial Intelligence (TeleAI), China Telecom, Beijing, China
Q
Qianrui Teng
Institute of Artificial Intelligence (TeleAI), China Telecom, Beijing, China
Y
Ying He
Institute of Artificial Intelligence (TeleAI), China Telecom, Beijing, China
Yuxuan Mu
Yuxuan Mu
Simon Fraser University
3D Computer VisionComputer Animation
C
Chao Ban
Institute of Artificial Intelligence (TeleAI), China Telecom, Beijing, China
Huayu Zhang
Huayu Zhang
Senior Engineer, Huawei Technologies Co., Ltd
Distributed SystemNetwork ScienceMachine LearningOptimizationGraph Theory
L
Lanxiang Zhou
Institute of Artificial Intelligence (TeleAI), China Telecom, Beijing, China
Z
Zerun Feng
Institute of Artificial Intelligence (TeleAI), China Telecom, Beijing, China
C
Chi Zhang
Institute of Artificial Intelligence (TeleAI), China Telecom, Beijing, China