Score
Designs, implements, or evaluates methods that compute numeric or categorical alignment scores between visual inputs and textual prompts — including prompt-conditioned scoring and representational similarity measures — to quantify how well images and language correspond. Builds analyses and tools to compare scores across label-consistent and inconsistent prompts, assess per-attribute presence fidelity, and flag likely generative hallucinations.
Text-to-image generation lacks a comprehensive evaluation framework that jointly assesses semantic alignment (prompt–image consistency) and visual fidelity. Method: We propose the first two-dimensional, structured taxonomy of quality metrics for this task, centered on “compositional quality” and “general quality.” Systematically analyzing over 30 metrics, we characterize their theoretical foundations, applicability domains, and associated benchmarks (e.g., COCO-TI, TIFA, Pick-a-Pic). Our analysis integrates literature review, taxonomic modeling, and empirical validation. Contribution/Results: We rigorously disentangle semantic alignment from visual fidelity, exposing critical limitations of existing metrics in fine-grained composition understanding, subjective perceptual modeling, and cross-domain generalization. The taxonomy fills a fundamental gap in evaluation paradigms, delivering a reproducible, extensible, and unified benchmarking framework. It provides both theoretical grounding and practical guidance for algorithm development and benchmark curation.
Automated evaluation metrics for text-to-image generation are often adopted empirically, lacking systematic validation against human judgments—particularly for compositional alignment involving objects, attributes, and relational semantics. Method: We propose a multidimensional analytical framework and conduct the first unified benchmarking of three metric families—VQA-based, embedding-based, and image-only—using large-scale human annotations as the gold standard. Contribution/Results: No single metric family dominates across all dimensions; VQA-based metrics are not universally superior, and certain embedding-based metrics exhibit higher discriminative power for fine-grained relational alignment. Image-only metrics show limited capacity to model compositional alignment. Our findings reveal strong task dependency in metric behavior, providing empirical guidance and methodological insights for principled metric selection in text-to-image evaluation.
This study addresses the limitations of existing text-to-image evaluation metrics, including poor generalization, a lack of fine-grained assessment, and misjudgments caused by reliance on fixed reference answers. To overcome these issues, this work proposes DEPICT, a training-free metric that discards static references in favor of a dynamic scoring mechanism based on expected consistency. Specifically, DEPICT leverages vision-language models to evaluate consistency through independent visual and textual question answering, employing multi-granularity weighted fusion to recover contextual information and effectively resolve the longstanding challenge of evaluating negation prompts. Extensive experiments across five benchmarks and eleven backbone networks demonstrate that DEPICT comprehensively outperforms existing training-free metrics and surpasses fine-tuned evaluators on two human-correlation benchmarks.
This work addresses the challenge that vision-language models struggle to accurately ground abstract semantics—such as idiomatic meanings of compound nouns—in high-fidelity image generation, where increased visual realism can interfere with compositional semantic understanding. To this end, the authors introduce the DIVA benchmark, which employs diagrammatic images to separately anchor literal and idiomatic interpretations. They further propose, for the first time, architecture-agnostic metrics: a semantic alignment gap (Δ) and a directional bias b(t), to quantify the disparity in visual grounding between these two semantic types. Experiments across eight state-of-the-art models reveal a pervasive literalness bias that persists despite model scaling and intensifies with higher visual fidelity, suggesting that iconographic abstraction enhances symbolic semantic alignment.
Text-to-image (T2I) models exhibit weak creativity, poor text–image alignment, and low persuasiveness when generating advertising images from implicit prompts. Method: We propose CAP—the first three-dimensional evaluation framework jointly assessing Creativity, Alignment (prompt fidelity), and Persuasiveness. CAP integrates multi-dimensional human evaluation, implicit-versus-explicit prompt comparison, quantitative measurement of visual-semantic consistency, and behavioral persuasion experiments to systematically uncover structural deficiencies of mainstream T2I models under implicit semantics. We further introduce a lightweight enhancement strategy targeting all three dimensions. Contribution/Results: CAP provides an interpretable, reproducible, and multi-objective evaluation and optimization paradigm for advertising image generation. Our enhancement strategy yields statistically significant average improvements of 18.7% across all three dimensions (p < 0.01), substantially elevating generation quality.
Systematic validation of prompt–image alignment evaluation in text-to-image (T2I) generation remains lacking; existing automatic metrics have not been rigorously assessed for quality, reliability, or cross-metric comparability against human judgments. Method: We propose the first comprehensive framework for alignment evaluation—introducing a skill-graded benchmark and a large-scale, multi-template human evaluation dataset with over 100K annotations; designing a skill-driven prompt taxonomy and a multi-round consistency scoring protocol; and developing QA-Metric, a question-answering–based automatic metric aligned with human judgment. Contribution/Results: Experiments demonstrate that QA-Metric significantly outperforms state-of-the-art methods on both our benchmark and TIFA160. Moreover, our analysis uncovers, for the first time, intrinsic connections among prompt ambiguity, model bias, and metric bias—revealing critical limitations in current alignment assessment paradigms.
This study investigates the feasibility of leveraging multimodal large language models (MLLMs) to automatically assess visual creativity in images without human annotations or model fine-tuning, while providing interpretable justifications. We present the first systematic evaluation of six MLLMs—Gemini, GPT, GLM, Kimi, Qwen, and Gemma—in a zero-shot setting for scoring the creativity of both AI-generated images and hand-drawn sketches. Experimental results demonstrate that model-generated scores exhibit significant correlation with human judgments, achieving a maximum Pearson correlation coefficient of 0.68. Furthermore, the models’ reasoning processes are interpretable, although this interpretability does not further improve scoring consistency with human assessments. Our findings highlight both the promising potential and inherent limitations of MLLMs for unsupervised evaluation of visual creativity.
This work addresses the high cost and subjectivity of human evaluation in assessing artistic creativity, as well as the limited interpretability of existing image-feature-based methods. The authors propose a multi-task fine-tuning framework leveraging Qwen2-VL-7B, which uniquely integrates a five-dimensional structured scoring rubric into system prompts. In a single forward pass, the model simultaneously generates highly accurate creativity scores—achieving a Pearson correlation coefficient exceeding 0.97 and a mean absolute error of approximately 3.95 on a 100-point scale—and produces aligned explanatory comments. Trained on a dataset of 1,000 paintings combining visual inputs, textual descriptions, and expert annotations, the model’s generated critiques achieve an SBERT similarity of 0.798 with human expert evaluations, substantially enhancing both the accuracy and interpretability of automated creativity assessment.
本文针对SVG生成缺乏领域特定评估协议的问题,提出了一种与人类判断对齐的评估框架,通过改进CLIP评分和训练视觉-语言模型来提高评估准确性。
This study addresses the tendency of small-scale open-source vision-language models (VLMs) to assign inflated, visually unsupported “flattering” scores when evaluating image–text alignment, thereby compromising assessment reliability. The authors construct a large-scale benchmark comprising 173,810 AI-generated fantasy character image–text pairs and introduce, for the first time, the “Bluffing Coefficient” to quantify the inconsistency between model-assigned scores and the visual evidence cited in their rationales. Combining multi-scale open-source VLMs (ranging from 450M to 8B parameters), automated evidence retrieval, and human verification, the work systematically demonstrates a strong negative correlation between model scale and flattering behavior (r = –0.96, p = 0.002): the smallest model (LFM2-VL) exhibits a flattering rate of 22.3%, while the largest (LLaVA-1.6) reduces it to just 6.0%.
This study addresses the prevalent misinterpretation of vision-text similarity in multimodal large language models (MLLMs) as evidence of cross-modal content interaction, which often obscures spurious alignment induced by shared language model pathways and weight anisotropy. Through controlled intervention experiments on thirteen mainstream MLLMs—incorporating Gaussian noise injection, CKA/SVCCA analysis, and principal angle cosine computation—this work proposes a novel metric termed the principal angle gap (PA gap). This metric effectively decouples weight-induced similarity from genuine multidimensional visual structure. Experimental results demonstrate that under hierarchical visual corruption, the PA gap tracks task accuracy more consistently than conventional scalar metrics. These findings confirm that internal alignment serves merely as a geometric diagnostic tool rather than a direct proxy for semantic content interaction.