unsupervised caption evaluation

Design and implement unsupervised automatic evaluation methods that score caption quality by measuring visual–textual alignment and content fidelity without labeled references; this includes building metrics and systems that verify whether caption mentions correspond to the depicted visual content and that combine multiple visual and textual alignment signals into a single quality score.

unsupervisedcaptionevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of reference-free image caption evaluation in capturing fine-grained semantic mismatches—such as hallucinations, missing attributes, or relational errors—by introducing a novel distributional scoring framework. For the first time, it models patch-level image and token-level text embeddings on the hypersphere as a multi-scale mixture of von Mises–Fisher distributions. The proposed method integrates bidirectional weighted KL divergence with global similarity to yield a comprehensive alignment score. It supports both single and multiple candidate captions and provides interpretable, decomposable diagnostic signals for local misalignments. Evaluated across multiple benchmarks, the approach achieves state-of-the-art correlation with human judgments while offering transparent, fine-grained error analysis capabilities.

distributional scoringfine-grained mismatchimage captioning

This work addresses the limitations of current image and video generation models, which suffer from poor text–image alignment due to low-quality captions generated by vision-language models—characterized by hallucinations, weak compositional reasoning, and insufficient fine-grained understanding. To overcome these issues, the authors propose a dual-path optimization strategy: first, constructing a high-quality, “clean” annotated dataset through hierarchical sampling that avoids reliance on web-scraped, copyrighted content; second, enhancing caption quality via context-aware alignment and supervised fine-tuning (SFT), augmented with a fine-tuned character detection model to produce structured captions that improve downstream usability. The study also introduces the first systematic categorization of caption evaluation metrics into generic and instance-anchored types. Experiments on open-source models demonstrate significant improvements in text–image alignment, particularly in caption accuracy and structural coherence.

caption qualityhallucinationimage-caption alignment

This study addresses the limitations of existing text-to-image evaluation metrics, including poor generalization, a lack of fine-grained assessment, and misjudgments caused by reliance on fixed reference answers. To overcome these issues, this work proposes DEPICT, a training-free metric that discards static references in favor of a dynamic scoring mechanism based on expected consistency. Specifically, DEPICT leverages vision-language models to evaluate consistency through independent visual and textual question answering, employing multi-granularity weighted fusion to recover contextual information and effectively resolve the longstanding challenge of evaluating negation prompts. Extensive experiments across five benchmarks and eleven backbone networks demonstrate that DEPICT comprehensively outperforms existing training-free metrics and surpasses fine-tuned evaluators on two human-correlation benchmarks.

Evaluation MetricsHallucination DetectionText-to-Image Alignment

This study addresses the unclear sensitivity of existing reference-free image-text evaluation metrics to semantically invariant perturbations. We present the first systematic assessment of five state-of-the-art evaluators under semantic-preserving transformations—including spatial manipulations, object scaling and category substitution, and neutral linguistic paraphrasing—and reveal that average score fluctuations of 6–9% can lead to ranking reversals in up to 37% of cases. To mitigate this instability, we propose a post-hoc calibration method that substantially enhances robustness to non-semantic variations while preserving high correlation with the original evaluator scores. Empirically, our approach reduces median absolute sensitivity by approximately 50%, significantly improving metric invariance without compromising alignment with human judgments.

caption evaluationimage-text alignmentmetric sensitivity

Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models

Jun 10, 2025
HZ
Huixuan Zhang
🏛️ Peking University | Wangxuan Institute of Computer Technology

Existing automatic evaluation methods for text-to-image alignment predominantly prioritize correlation with human judgments while neglecting foundational trustworthiness attributes—consistency and robustness—essential for reliable assessment. Method: The authors formally define and empirically validate these two trustworthiness properties through systematic, controlled experiments across diverse diffusion models (e.g., Stable Diffusion, SDXL, DALL·E 3) and alignment metrics (e.g., CLIPScore, TIFA, Pick-a-Pic), complemented by attribution analysis. Contribution/Results: All 12 mainstream evaluation methods violate at least one trustworthiness property. To address this, the authors propose a reproducible and scalable framework for evaluation improvement—already adopted by three top-tier conference papers—thereby shifting the paradigm of text–image alignment evaluation from “correlation-oriented” to “trustworthiness-oriented.”

Evaluating image-text alignment in text-to-image modelsIdentifying gaps in current evaluation frameworksProposing improvements for reliable alignment assessment

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing image captioning evaluation methods, which predominantly rely on human-annotated reference descriptions and struggle to assess semantic faithfulness. To overcome this, the authors propose a reference-free evaluation framework that judges caption quality through semantic-equivalent reconstruction: captions are used to reconstruct images, and the reconstructed images are evaluated based on their performance consistency with the original images in downstream vision-language tasks. Departing from pixel-level reconstruction, this approach introduces a novel evaluation principle centered on semantic equivalence and incorporates a task-conditioned scoring mechanism. Leveraging a newly constructed Captioning Turing Test Dataset (CTTD), the study establishes the first reference-free evaluation system that effectively captures the semantic fidelity of captions while substantially reducing annotation costs.

caption evaluationimage captioningreference-free

This work addresses the limitations of existing reference-dependent evaluation metrics for remote sensing image captioning, which are prone to annotator-style biases that obscure the true capabilities of multimodal large language models (MLLMs) and mislead judgments about the necessity of fine-tuning. To overcome this, the authors propose ReconScore, a reference-free metric that assesses semantic fidelity by reconstructing the original visual content from generated captions. Building upon ReconScore, they introduce RemoteDescriber—a training-free method incorporating an iterative self-correction mechanism to enhance caption accuracy. Experiments demonstrate that off-the-shelf MLLMs outperform fine-tuned counterparts in genuine zero-shot settings, and RemoteDescriber achieves state-of-the-art performance across three remote sensing captioning benchmarks. Furthermore, ReconScore is shown to be more reliable and equitable than conventional evaluation metrics.

evaluation biasfoundation modelsreference-free evaluation

This work addresses the absence of text rendering quality assessment methods aligned with human perception in current text-to-image generation models, as mainstream OCR and vision-language models struggle to accurately capture visual artifacts. The paper introduces the Textual Image Quality Assessment (TIQA) task, which quantifies the fidelity of rendered text in generated images by predicting a scalar score aligned with human mean opinion scores (MOS). To support this task, two MOS-annotated datasets are constructed, and a lightweight, no-reference evaluation model, ANTIQa, is proposed. By incorporating text-specific biases, ANTIQa improves PLCC correlation by at least 0.05 on TIQA-Crops and 0.08 on TIQA-Images. When applied to rerank generated outputs, it increases average human-rated quality by 14%.

evaluation metricshuman perceptionquality assessment

This work proposes a comparative learning approach based on human relative preference judgments to address the high cost and subjectivity of manual scoring in image–text matching quality assessment. By replacing conventional regression modeling with pairwise comparison learning, the method significantly reduces annotation costs while improving inter-annotator consistency. The model leverages ResNet-50 for visual feature extraction and MiniLM for textual features within a paired comparison framework. Experimental results demonstrate that performance steadily improves with increasing data volume, achieving a Pearson correlation coefficient of 0.7609—approaching the performance of regression-based models. Human evaluations further confirm that the proposed approach yields higher annotation efficiency and stronger consistency compared to traditional scoring methods.

caption evaluationcomparative judgmentshuman annotation

Existing text evaluation metrics, despite exhibiting high correlation with human judgments, are vulnerable to strategic manipulation and lack robustness against irrelevant perturbations. This work introduces dual criteria—statistical alignment and strategic alignment—and formally defines strategic alignment for the first time, establishing a principled framework that encompasses human correlation, degradation sensitivity, and robustness to manipulation. Grounded in mutual information theory, the authors propose a unified metric design framework composed of four components: information measures, estimation methods, text representations, and prediction mechanisms. Empirical results demonstrate that strong correlation with human scores does not imply strategic robustness; the proposed metrics significantly enhance resistance to manipulation across peer review, summarization, and question-answering tasks while maintaining high agreement with human evaluations.

evaluation metricsnatural language generationreference-based scoring

Hot Scholars

HL

Hyeongkeun Lee

KAIST
Deep LearningMultimodal LearningVideo Understanding
SK

Sungchul Kim

Adobe
Data miningMachine learningBioinformatics
RA

Ryan A. Rossi

Adobe Research
Machine LearningPersonalizationGraph Representation LearningGraph ML
TK

Taewhan Kim

Seoul National University, Department of Electrical and Computer Engineering
Electronic Design Automation
CY

Chieh-Yang Huang

MetaMetrics Inc.
Natural Language ProcessingDeep LearningHCICrowdsourcing