🤖 AI Summary
This work addresses instruction-based image editing (IIE)—enabling precise and controllable image manipulation through natural language commands. To advance this emerging field, we establish a unified conceptual framework encompassing task formulation, data curation, model architectures, evaluation protocols, and real-world applicability. We further introduce CDD-IIE Bench, the first comprehensive benchmark for IIE, which facilitates multi-dimensional and fine-grained performance diagnosis. By systematically integrating techniques from GANs, diffusion models, autoregressive models, and large language/vision-language models, we conduct an extensive empirical comparison of prominent open-source methods, elucidating their respective strengths and limitations. Our analysis clarifies key evolutionary trajectories in IIE methodologies and provides the community with a standardized evaluation toolkit and actionable directions for future research.
📝 Abstract
Instruction-based Image Editing (IIE) aims to transform a given image into a new one based on textual instructions. Advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) have accelerated progress toward practical ``one-sentence image editing" systems. This survey presents a systematic taxonomy and comprehensive review of IIE research, structured around five core dimensions: (1) task definition and hierarchical categorization of editing operations, (2) methodologies for training data construction, (3) architectural evolution from GAN-based to diffusion and autoregressive paradigms, (4) standardized evaluation metrics and benchmark development, and (5) introduction of commercial solutions. Our analysis shows critical technological milestones across model generations. We further propose a Comprehensive, in-Depth, and Diagnostic benchmark for IIE task (CDD-IIE Bench), which can rigorously assess the multiple aspects of model performance. Through empirical comparisons of open-source solutions, we highlight their respective capabilities and limitations. Finally, we discuss future research directions to advance the field.