Score
Design and implement unsupervised automatic evaluation methods that score caption quality by measuring visual–textual alignment and content fidelity without labeled references; this includes building metrics and systems that verify whether caption mentions correspond to the depicted visual content and that combine multiple visual and textual alignment signals into a single quality score.
This study addresses the weak correlation between automated image captioning evaluation metrics and human judgments. We systematically survey over 70 existing metrics and propose, for the first time, a comprehensive, hierarchically structured taxonomy. Empirical analysis reveals that five widely adopted metrics—including BLEU and METEOR—exhibit consistently low Spearman correlations (<0.3) with human ratings across diverse benchmarks. To overcome this limitation, we introduce EnsembEval, a linear regression-based ensemble framework that fuses multiple metrics. Trained on a single dataset, EnsembEval achieves statistically significant improvements in both Spearman and Pearson correlations (average gain >0.15) across five out-of-domain test sets, demonstrating strong generalization. Our work contributes both an interpretable, principled taxonomy for caption evaluation metrics and a reusable, effective ensemble methodology—advancing the reliability and applicability of automatic image caption assessment.
This work addresses the challenge of reference-free image caption evaluation in capturing fine-grained semantic mismatches—such as hallucinations, missing attributes, or relational errors—by introducing a novel distributional scoring framework. For the first time, it models patch-level image and token-level text embeddings on the hypersphere as a multi-scale mixture of von Mises–Fisher distributions. The proposed method integrates bidirectional weighted KL divergence with global similarity to yield a comprehensive alignment score. It supports both single and multiple candidate captions and provides interpretable, decomposable diagnostic signals for local misalignments. Evaluated across multiple benchmarks, the approach achieves state-of-the-art correlation with human judgments while offering transparent, fine-grained error analysis capabilities.
This work addresses the limitations of current image and video generation models, which suffer from poor text–image alignment due to low-quality captions generated by vision-language models—characterized by hallucinations, weak compositional reasoning, and insufficient fine-grained understanding. To overcome these issues, the authors propose a dual-path optimization strategy: first, constructing a high-quality, “clean” annotated dataset through hierarchical sampling that avoids reliance on web-scraped, copyrighted content; second, enhancing caption quality via context-aware alignment and supervised fine-tuning (SFT), augmented with a fine-tuned character detection model to produce structured captions that improve downstream usability. The study also introduces the first systematic categorization of caption evaluation metrics into generic and instance-anchored types. Experiments on open-source models demonstrate significant improvements in text–image alignment, particularly in caption accuracy and structural coherence.
This study addresses the limitations of existing text-to-image evaluation metrics, including poor generalization, a lack of fine-grained assessment, and misjudgments caused by reliance on fixed reference answers. To overcome these issues, this work proposes DEPICT, a training-free metric that discards static references in favor of a dynamic scoring mechanism based on expected consistency. Specifically, DEPICT leverages vision-language models to evaluate consistency through independent visual and textual question answering, employing multi-granularity weighted fusion to recover contextual information and effectively resolve the longstanding challenge of evaluating negation prompts. Extensive experiments across five benchmarks and eleven backbone networks demonstrate that DEPICT comprehensively outperforms existing training-free metrics and surpasses fine-tuned evaluators on two human-correlation benchmarks.
This study addresses the unclear sensitivity of existing reference-free image-text evaluation metrics to semantically invariant perturbations. We present the first systematic assessment of five state-of-the-art evaluators under semantic-preserving transformations—including spatial manipulations, object scaling and category substitution, and neutral linguistic paraphrasing—and reveal that average score fluctuations of 6–9% can lead to ranking reversals in up to 37% of cases. To mitigate this instability, we propose a post-hoc calibration method that substantially enhances robustness to non-semantic variations while preserving high correlation with the original evaluator scores. Empirically, our approach reduces median absolute sensitivity by approximately 50%, significantly improving metric invariance without compromising alignment with human judgments.
Existing automatic evaluation methods for text-to-image alignment predominantly prioritize correlation with human judgments while neglecting foundational trustworthiness attributes—consistency and robustness—essential for reliable assessment. Method: The authors formally define and empirically validate these two trustworthiness properties through systematic, controlled experiments across diverse diffusion models (e.g., Stable Diffusion, SDXL, DALL·E 3) and alignment metrics (e.g., CLIPScore, TIFA, Pick-a-Pic), complemented by attribution analysis. Contribution/Results: All 12 mainstream evaluation methods violate at least one trustworthiness property. To address this, the authors propose a reproducible and scalable framework for evaluation improvement—already adopted by three top-tier conference papers—thereby shifting the paradigm of text–image alignment evaluation from “correlation-oriented” to “trustworthiness-oriented.”
This work addresses the limitations of existing image captioning evaluation methods, which predominantly rely on human-annotated reference descriptions and struggle to assess semantic faithfulness. To overcome this, the authors propose a reference-free evaluation framework that judges caption quality through semantic-equivalent reconstruction: captions are used to reconstruct images, and the reconstructed images are evaluated based on their performance consistency with the original images in downstream vision-language tasks. Departing from pixel-level reconstruction, this approach introduces a novel evaluation principle centered on semantic equivalence and incorporates a task-conditioned scoring mechanism. Leveraging a newly constructed Captioning Turing Test Dataset (CTTD), the study establishes the first reference-free evaluation system that effectively captures the semantic fidelity of captions while substantially reducing annotation costs.
This work addresses the limitations of existing reference-dependent evaluation metrics for remote sensing image captioning, which are prone to annotator-style biases that obscure the true capabilities of multimodal large language models (MLLMs) and mislead judgments about the necessity of fine-tuning. To overcome this, the authors propose ReconScore, a reference-free metric that assesses semantic fidelity by reconstructing the original visual content from generated captions. Building upon ReconScore, they introduce RemoteDescriber—a training-free method incorporating an iterative self-correction mechanism to enhance caption accuracy. Experiments demonstrate that off-the-shelf MLLMs outperform fine-tuned counterparts in genuine zero-shot settings, and RemoteDescriber achieves state-of-the-art performance across three remote sensing captioning benchmarks. Furthermore, ReconScore is shown to be more reliable and equitable than conventional evaluation metrics.
This work addresses the absence of text rendering quality assessment methods aligned with human perception in current text-to-image generation models, as mainstream OCR and vision-language models struggle to accurately capture visual artifacts. The paper introduces the Textual Image Quality Assessment (TIQA) task, which quantifies the fidelity of rendered text in generated images by predicting a scalar score aligned with human mean opinion scores (MOS). To support this task, two MOS-annotated datasets are constructed, and a lightweight, no-reference evaluation model, ANTIQa, is proposed. By incorporating text-specific biases, ANTIQa improves PLCC correlation by at least 0.05 on TIQA-Crops and 0.08 on TIQA-Images. When applied to rerank generated outputs, it increases average human-rated quality by 14%.
This work proposes a comparative learning approach based on human relative preference judgments to address the high cost and subjectivity of manual scoring in image–text matching quality assessment. By replacing conventional regression modeling with pairwise comparison learning, the method significantly reduces annotation costs while improving inter-annotator consistency. The model leverages ResNet-50 for visual feature extraction and MiniLM for textual features within a paired comparison framework. Experimental results demonstrate that performance steadily improves with increasing data volume, achieving a Pearson correlation coefficient of 0.7609—approaching the performance of regression-based models. Human evaluations further confirm that the proposed approach yields higher annotation efficiency and stronger consistency compared to traditional scoring methods.
Existing text evaluation metrics, despite exhibiting high correlation with human judgments, are vulnerable to strategic manipulation and lack robustness against irrelevant perturbations. This work introduces dual criteria—statistical alignment and strategic alignment—and formally defines strategic alignment for the first time, establishing a principled framework that encompasses human correlation, degradation sensitivity, and robustness to manipulation. Grounded in mutual information theory, the authors propose a unified metric design framework composed of four components: information measures, estimation methods, text representations, and prediction mechanisms. Empirical results demonstrate that strong correlation with human scores does not imply strategic robustness; the proposed metrics significantly enhance resistance to manipulation across peer review, summarization, and question-answering tasks while maintaining high agreement with human evaluations.