MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of evaluation benchmarks aligned with human preferences for multi-image editing (MIE) tasks, as existing methods primarily focus on single-image editing. To bridge this gap, the authors introduce MIE-Bench, the first large-scale benchmark for MIE, encompassing 16 task categories, 3,000 instances, and 108K human ratings. They further propose MIEScore, an automatic evaluation method based on multimodal large language models, which achieves precise alignment with human judgments through multidimensional supervised fine-tuning. Experimental results demonstrate that MIEScore attains state-of-the-art performance in human preference alignment on MIE-Bench and exhibits strong generalization capabilities across other image editing evaluation datasets.
📝 Abstract
Recent advances in unified multimodal models have significantly improved text-guided image editing abilities. In particular, models such as Nano-Banana-Pro and GPT-Image-2 demonstrate emerging capabilities in multi-source image editing (MIE), including tasks such as object synthesis, person-background composition, and cross-image style fusion. However, existing benchmarks and image editing assessment (IEQA) methods remain primarily focused on single-image editing tasks and largely overlook the more challenging setting of MIE. This highlights the urgent need for a comprehensive and human-aligned benchmark for MIE. To this end, we introduce MIE-Bench, the first large-scale multiple image editing benchmark with fine-grained human preference annotations. Specifically, MIE-Bench includes 3,000 editing instances across 16 tasks, each involving more than two source images and an editing prompt, together with 36K edited images produced by 12 state-of-the-art editing models and over 108K mean opinion scores (MOSs) covering visual quality, instruction following, and attribute preservation. Based on MIE-Bench, we propose MIEScore, a multimodal large language model (MLLM)-based evaluation model enhanced with skill optimization and multi-dimensional supervised fine-tuning, to provide human-aligned feedback for MIE. Extensive experiments show that MIEScore achieves state-of-the-art performance in aligning with human preferences and generalizes well across other IEQA datasets. Both the dataset and the model are available at https://github.com/IntMeGroup/MIEScore.
Problem

Research questions and friction points this paper is trying to address.

multi-source image editing
image editing assessment
human-aligned evaluation
benchmark
MIE
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-source image editing
human-aligned evaluation
MIE-Bench
MIEScore
multimodal large language model