Evaluating the Evaluators: Diagnosing Large Multimodal Models for AI-Generated Image Assessment

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of conventional metrics in capturing fine-grained alignment and generative artifacts, as well as the absence of unified benchmarks evaluating semantics, quality, authenticity, and responsibility. To bridge this gap, we propose SQUARE-Bench, a novel benchmark featuring a pioneering four-in-one evaluation framework that uniquely incorporates the responsibility dimension. It establishes a fine-grained taxonomy comprising 38 sub-dimensions across nearly 10,000 real and generated images, and introduces a large multimodal model (LMM)-guided iterative editing approach. Our findings reveal that top-tier proprietary models have surpassed human expert baselines, albeit with substantial performance disparities. Furthermore, LMM-guided editing selectively enhances image performance in terms of semantics, authenticity, and responsibility.
📝 Abstract
With the rapid advancement of text-to-image (T2I) generation, robust evaluation becomes critical yet challenging, as traditional metrics fail to capture fine-grained alignment and generative artifacts. While large multimodal models (LMMs) are increasingly adopted as evaluators, existing benchmarks typically study semantic understanding, quality perception, and authenticity identification in isolation, while largely neglecting responsibility detection. This leaves a gap in unified and comprehensive validation. To bridge this gap, we introduce SQUARE-Bench, a comprehensive benchmark that systematically evaluates LMM capabilities as evaluators of AI-generated images across four aspects: Semantics, Quality, Authenticity, and Responsibility. SQUARE-Bench introduces a granular taxonomy of 38 sub-dimensions to evaluate nearly 10K AI-generated images sampled from 22 diverse models, ranging from legacy to state-of-the-art generators, complemented by over 3K real-world images. The images are annotated with curated question-answering pairs. Extensive experiments on 23 LMMs reveal that top proprietary models, such as Gemini-3-Pro, already outperform the individual human expert baseline. However, the performance gap between models remains significant, exhibiting notable disparities in fine-grained inference and domain-specific robustness. Beyond benchmarking, we conduct a proof-of-concept study of LMM-guided iterative editing, in which dimension-specific LMMs provide diagnostic feedback to fixed image editors. The resulting guided system yields selective improvements in semantics, authenticity, and responsibility, while exhibiting a consistent visual-quality trade-off. SQUARE-Bench can serve as both a diagnostic tool for characterizing LMM evaluator capabilities and studying their use in T2I generation refinement. The benchmark and dataset will be released upon publication.
Problem

Research questions and friction points this paper is trying to address.

AI-generated image assessment
large multimodal models
evaluation benchmark
text-to-image generation
responsibility detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Multimodal Models
AI-Generated Image Assessment
Benchmark
Fine-grained Taxonomy
Iterative Editing
🔎 Similar Papers