🤖 AI Summary
This study addresses the limitation of existing image assessment models that output only scalar scores without defect localization or interpretability. We propose a unified image assessment framework utilizing a text-native grid representation as its interface to jointly predict quality scores, localize defective regions, and generate interpretable descriptions within a single forward pass. Methodologically, we integrate supervised fine-tuning with Group Relative Policy Optimization (GRPO) reinforcement learning—leveraging Dice and format-based rewards—to achieve post-training optimization under heterogeneous spatial supervision, alongside introducing a parameter-free parser for structured prediction conversion. Experimental results demonstrate that our model attains an overall scoring SRCC of 0.601, outperforming the Gemini baseline, while achieving superior defect localization performance over both general-purpose and specialized models across most benchmarks.
📝 Abstract
Existing synthetic image evaluators typically provide only a scalar quality score and do not identify the image regions that support it. We introduce VIEScore2, a unified evaluator for image generation and editing tasks with optional conditioning images. VIEScore2 represents an image as an N x N grid and jointly predicts quality scores and defect locations in a single model pass. Its text-native grid representation provides a common interface for heterogeneous spatial supervision and enables directly verifiable post-training objectives. We train on 38K examples spanning score-only, localization-only, and joint supervision across generation and editing tasks. Starting from supervised fine-tuning, we further apply GRPO to improve defect localization using rewards that combine cell-level Dice overlap, score accuracy, and output-format validity. A parameter-free parser converts the structured predictions into readable explanations. On the primary suite, VIEScore2 achieves an overall-score SRCC of 0.601, compared with 0.491 for Gemini-3-Flash, the strongest zero-shot general-purpose VLM baseline under matched inputs. For defect localization, VIEScore2 outperforms both general-purpose VLMs and specialized spatial evaluators on three of six benchmarks in per-image grid IoU and ranks among the top three on five, including datasets beyond its training sources.