semantic similarity evaluation

Designing and applying perceptual and semantic similarity metrics (e.g., cosine/perceptual scores) to measure how well models preserve identity, intent, or visual similarity while enabling desired variation or edits, and to benchmark fidelity and temporal/semantic coherence.

semanticsimilarityevaluation

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitation of existing image perceptual similarity metrics in capturing the context-dependent nature of human judgments across diverse semantic dimensions—such as shape or color. To this end, the authors construct a large-scale dataset of human triplet similarity judgments annotated with free-form semantic dimensions and leverage it to fine-tune state-of-the-art vision-language models. They propose TPIPS (Text-Prompted Perceptual Image Similarity), a novel metric that enables dynamic specification of similarity semantics via natural language prompts, thereby overcoming the rigidity of conventional single-purpose similarity measures. Experiments demonstrate that TPIPS significantly outperforms existing methods in aligning with human judgments and generalizing across data distributions. The approach also proves effective in downstream applications including text-guided image retrieval, compositional search, and fine-grained evaluation of generative models.

context-dependent judgmentimage tripletsperceptual metric

Measuring and predicting visual fidelity

Jul 15, 2025
BW
Benjamin Watson
🏛️ Northwestern University | University of Alberta

Accurately measuring visual fidelity and predicting human perception remains challenging, particularly in distinguishing perceptual differences across object categories and simplification types. Method: We propose a systematic evaluation framework integrating polygon mesh simplification (using two distinct algorithms) with psychophysical experiments—namely naming time, subjective rating, and pairwise preference judgments—to rigorously compare perceptual responses to animals versus man-made objects under controlled simplification levels. Contribution/Results: This work introduces the first three-factor controlled analysis jointly varying simplification type, degree, and object category. We find naming time and preference tasks exhibit higher sensitivity to fidelity degradation than subjective ratings; conversely, state-of-the-art image- and mesh-based automatic metrics reliably predict only subjective ratings, failing to model higher-order perceptual decisions. These findings expose fundamental cognitive limitations of prevailing fidelity metrics and provide empirical grounding and a methodological paradigm for developing human-centered graphics fidelity assessment frameworks.

Comparing experimental techniques for fidelity assessmentEvaluating automatic prediction of visual fidelity measuresMeasuring visual fidelity of polygonal models

How Small Transformation Expose the Weakness of Semantic Similarity Measures

Sep 08, 2025
SL
Serge Lionel Nikiema
🏛️ University of Luxembourg

This study addresses the fundamental question of whether semantic similarity measures genuinely comprehend semantic relationships. We propose the first evaluation framework based on controlled, small-scale semantic transformations to systematically assess the semantic discrimination capability of 18 state-of-the-art methods—including bag-of-words, embedding-based, LLM-based, and structure-aware models—on software engineering texts and code. Experiments reveal that mainstream embedding methods exhibit up to 99.9% misclassification rates in semantic opposition scenarios, exposing their reliance on superficial surface patterns. Substituting cosine similarity for Euclidean distance improves performance by 24–66%. LLM-based methods demonstrate superior fine-grained semantic distinction. Critically, our framework uncovers a foundational limitation in existing measures: their failure to capture semantic essence. It establishes the first reproducible, scalable benchmark paradigm for trustworthy semantic computation in software engineering contexts.

Evaluating semantic similarity measures for software engineering tasksIdentifying flaws where methods confuse opposites and synonymsTesting 18 methods including embeddings and LLMs on semantic understanding

Foundation Models Boost Low-Level Perceptual Similarity Metrics

Sep 11, 2024
AG
Abhijay Ghildyal
🏛️ Portland State University | Sony Interactive Entertainment

To address the misalignment between model predictions and human visual perception in full-reference image quality assessment (FR-IQA), this paper proposes a zero-shot, fine-tuning-free feature distance metric. Instead of relying solely on final-layer outputs or embedding vectors, the method systematically exploits intermediate-layer features from pre-trained vision foundation models (e.g., ViT or CNN). Quality scores are computed directly via parameter-free distances—Euclidean distance or cosine similarity—between corresponding intermediate feature maps. This work is the first to empirically demonstrate that intermediate-layer features strongly encode low-level perceptual similarity, challenging the conventional FR-IQA paradigm that depends either on end-to-end learning or high-level semantic features. Evaluated on multiple standard benchmarks, the proposed method achieves state-of-the-art performance without any training—outperforming classical metrics (PSNR, SSIM) and recent learning-based approaches (e.g., DISTS, PieAPP).

Human Perception AlignmentImage Quality AssessmentIntermediate Layer Representation

MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences

Oct 03, 2024
GI
Genta Indra Winata
🏛️ Capital One | University of Toronto | Monash University Indonesia | Boston University

To address the misalignment between automatic evaluation metrics and human preferences in generative tasks, this paper proposes MetaMetrics—a calibratable meta-metric that supervisely weights and fuses existing metrics to model fine-grained human preferences across multimodal (language/vision), multilingual, and multi-domain settings. Methodologically, it introduces the first preference-dimension-aware metric calibration framework, enabling cross-modal unified evaluation and plug-and-play integration. The approach combines supervised meta-learning, multi-task joint optimization, and explicit modeling of human preference annotations. Experiments demonstrate that MetaMetrics significantly improves correlation with human judgments across multilingual text and vision generation tasks (average Kendall’s τ increase of +18.7%). Moreover, it exhibits strong generalization to unseen domains and models, maintaining robust alignment with human preferences without task-specific retraining.

Calibrate metrics to align with human preferences.Evaluate generation tasks across different modalities.Optimize existing metrics for multilingual and multi-domain scenarios.

Latest Papers

What's happening recently
View more

Existing image quality assessment methods primarily emphasize visual fidelity and struggle to capture the preservation of semantic content in low-level image processing. This work formally introduces the task of “semantic similarity” evaluation and proposes a structured semantic representation framework that decouples foreground and background entities. By integrating open-world category and relational modeling, the framework constructs image triplets representing semantic structures and introduces a Triplet Semantic Similarity (T3S) score for quantitative assessment. Experiments on COCO and SPA-Data demonstrate that T3S significantly outperforms existing fidelity-based metrics and semantic-level baselines, offering a more accurate characterization of progressive semantic changes under various image degradations.

Image Quality AssessmentLow-Level Image ProcessingSemantic Content Preservation

Existing embedding-based approaches to analyzing creative processes rely on static semantic similarity, which struggles to capture pivotal transitions in creative trajectories and lacks cross-domain comparability. This work systematically identifies, for the first time, three open challenges inherent in applying embedding methods to creativity analysis. To address these limitations, the paper proposes a novel paradigm that integrates large language models into a context-aware intervention framework for dynamically parsing multimodal design trajectories. By enhancing sensitivity to conversational context, the approach significantly improves the segmentation, representation, and evaluation of creative processes. This advancement lays both a theoretical foundation and a technical pathway toward building analytical frameworks with greater creative sensitivity.

AI creativity support toolscreative processcross-domain comparison

Existing text-to-image generation evaluation metrics struggle to distinguish between foreground subjects and background due to their global image processing, leading to inaccurate assessments of concept fidelity and prompt adherence. This work proposes MaSC, the first spatially decomposed evaluation paradigm that leverages external foreground masks to decouple assessment into subject-specific concept fidelity and background-related prompt following. Built upon a frozen SigLIP2 SO400M-NaFlex model, MaSC incorporates masked maximum cosine matching, background-pooled embeddings, and subject-removed prompt contrastive scoring. Experiments demonstrate that MaSC achieves a Krippendorff’s α of 0.471 for concept fidelity on DreamBench++ and an identity recognition AUC of 0.992 on ORIDa, significantly outperforming CLIP-T baselines and exhibiting stronger alignment with human perception.

concept preservationevaluation metricmasked similarity

Existing image similarity metrics such as LPIPS and CLIP often fail to align with human subjective judgments in text-to-image generation tasks, particularly in personalized or context-sensitive scenarios. This work proposes CLPIPS, which uniquely leverages user-provided ranking feedback on generated images to fine-tune the layer combination weights of LPIPS through a lightweight adaptation. By optimizing these weights using a margin-based ranking loss on human-annotated data, CLPIPS achieves personalized alignment with perceptual similarity. Consistency with human judgments is evaluated using Spearman’s rank correlation coefficient and intraclass correlation coefficient. Experimental results demonstrate that CLPIPS significantly outperforms the original LPIPS in capturing user preferences, thereby validating the efficacy of lightweight, personalized fine-tuning for perceptual similarity assessment.

human judgment alignmenthuman-in-the-loopimage similarity metrics

This study addresses the misalignment between existing graph similarity metrics and human visual perception, which undermines the effectiveness of visualization recommendation systems. Through a three-stage human-subject experiment, the authors systematically collected similarity judgments for 1,881 node-link diagrams, establishing the first human-grounded benchmark for graph perception. They evaluated the alignment of 16 conventional similarity measures and multimodal large language models (MLLMs) with this benchmark. Findings reveal that humans prioritize global shape and edge density when judging graph similarity. Among traditional metrics, Portrait divergence performs best but still shows limited alignment. In contrast, GPT-5 substantially outperforms conventional methods while offering interpretability, and Claude Sonnet 4.5 achieves the highest computational efficiency. The results demonstrate that MLLMs can serve as highly aligned, interpretable proxies for human perceptual judgment in graph similarity assessment.

computational measuresgraph similarityhuman perception

Hot Scholars

YH

Yupeng Hu

Shandong University
Multimedia Information RetrievalData Mining and Knowledge Discovery
MK

Meenakshi Khosla

UC San Diego
Computational NeuroscienceArtificial IntelligenceVisionAudition
TP

Themis Palpanas

Distinguished Professor, University Paris Cite, French University Institute (IUF)
data managementdata sciencedata/time seriesanomaly detection
GT

Giorgos Tolias

Czech Technical University in Prague
Computer VisionImage retrieval
XS

Xuemeng Song

City University of Hong Kong
Information RetrievalMultimedia Analysis