measure perceptual similarity

Designs, implements, and evaluates quantitative metrics and algorithms that estimate perceptual similarity between data items or their representations, including distance/similarity scores (e.g., cosine, Jaccard), structural-similarity measures, weighting schemes, and ranking procedures. Builds statistical tests and validation procedures for those measures and tools for computing nearest-neighbor rankings, redundancy detection, and feature-wise similarity assessments.

measureperceptualsimilarity

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.42
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$168K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Conventional dimensionality reduction (DR) evaluation suffers from systematic bias due to the frequent adoption of highly correlated metrics, leading to overemphasis on specific structural properties. Method: We propose an empirically grounded metric redundancy reduction framework: first computing Pearson correlation matrices across diverse datasets and DR algorithms; then applying clustering to identify functionally redundant metric groups; and finally retaining only the most representative metric per group—replacing subjective, intent-driven metric selection with objective, behavior-based clustering. Contribution/Results: Our approach significantly improves cross-dataset and cross-algorithm stability of DR evaluations, effectively mitigating structural biases inherent in traditional assessment protocols. Experimental validation demonstrates enhanced reproducibility and generalizability, establishing a principled, data-driven framework for fair and robust comparative evaluation of DR methods.

Clustering metrics by empirical correlations to avoid overlapImproving stability and reliability of DR projection evaluationsReducing bias in dimensionality reduction evaluation metrics selection

This study addresses the lack of systematic and neutral comparisons among similarity measures for categorical datasets. It presents the first comprehensive evaluation of several prominent methods—including edge-count tests, constrained minimum distance, graph-based tests, Classifier Two-Sample Tests (C2ST), and the Maximum Mean Discrepancy with Categorical Metrics (MMCM)—assessing their ability to detect distributional differences and their computational costs in both two-sample and multi-sample settings. The results demonstrate that the Friedman–Rafsky test achieves the best overall performance in two-sample tasks, while MMCM excels in multi-sample scenarios by offering both high statistical power and computational efficiency. This work provides empirical evidence and practical guidance for selecting appropriate similarity measures when analyzing categorical data.

categorical datadataset similaritymulti-sample comparison

This work proposes a novel, interpretable perceptual similarity metric that jointly models texture and chromatic distortions—addressing the limited perceptual consistency of existing image similarity measures under complex distortions involving both texture and color. The method quantifies texture differences using Earth Mover’s Distance and computes chromatic discrepancies in the perceptually uniform Oklab color space. By integrating these components, the approach not only achieves superior perceptual alignment but also offers visual interpretability to enhance evaluation transparency. Evaluated on the Berkeley-Adobe dataset featuring non-traditional distortions, the proposed metric significantly outperforms state-of-the-art methods, demonstrating particularly robust performance and higher perceptual consistency in scenarios involving shape-related distortions.

chromatic assessmentexplainabilityimage similarity

Traditional spreadsheet template identification suffers from poor distinguishability due to high similarity in both spatial layout and data-type patterns. To address this, we propose a fine-grained similarity metric that jointly encodes semantic embeddings, data-type representations, and cell-level spatial coordinates. Our method is the first to integrate Chamfer and Hausdorff distances in an unsupervised framework, enabling holistic modeling of semantic, typological, and geometric information. Operating at the cell level, it achieves a perfect Adjusted Rand Index of 1.00 on the FUSTE benchmark—significantly outperforming the graph-based baseline Mondrian (0.90)—and enables exact template clustering and reconstruction. The approach supports downstream applications including retrieval-augmented generation and large-scale data cleaning. By delivering scalable, high-precision template discovery, it establishes a new paradigm for structured spreadsheet analysis.

Enabling automated template discovery for large-scale spreadsheet collectionsOvercoming limitations of traditional structural similarity detection methodsQuantifying spreadsheet similarity using semantic embeddings and spatial layouts

Methods for quantifying dataset similarity: a review, taxonomy and comparison

Dec 07, 2023
MS
Marieke Stolte
🏛️ TU Dortmund University

This study addresses the critical need for quantifying dataset similarity in model generalization, transfer learning, simulation calibration, and two-sample testing. We systematically survey 118 similarity quantification methods and propose the first ten-dimensional classification framework, organizing approaches into seven technical categories: statistical distances (e.g., Wasserstein distance, Maximum Mean Discrepancy), kernel-based methods, information-theoretic measures, dimensionality-reduction embeddings, permutation tests, generative-model-based discriminators, and Gaussian process likelihood ratios. We develop a multi-dimensional evaluation system balancing theoretical guarantees, interpretability, and practical applicability, yielding a structured recommendation matrix aligned with task requirements and data characteristics. Furthermore, we introduce the first open-source, interactive tool for method selection—enabling real-time filtering and parameter configuration—to significantly enhance both selection efficiency and deployment suitability.

Compare 118 methods on applicability and interpretabilityProvide recommendations for selecting dataset similarity measuresReview and classify methods for quantifying dataset similarity

Latest Papers

What's happening recently
View more

This work proposes a language- and syntax-agnostic approach to string similarity measurement by introducing co-occurrence matrices (COM) and run-length matrices (RLM)—concepts originally from image texture analysis—into the domain of string representation to construct purely statistical, language-independent features. The method integrates multiple statistical measures, including COM, RLM, longest common subsequence, and edit distance. Evaluated on synthetic datasets, COM and RLM significantly outperformed baseline methods in three out of four experiments (p < 0.001). In real-world text plagiarism detection tasks, RLM achieved the best performance, demonstrating the effectiveness and generalizability of the proposed statistical features.

language-independentplagiarism detectionstatistical features

This study addresses the lack of systematic and impartial evaluation of existing methods for measuring distributional similarity in numerical data, which hinders informed selection in practice. The authors construct the first comprehensive benchmarking framework encompassing 36 similarity measures for continuous data—including statistical tests, distance-based metrics, and embedding approaches—and evaluate their discriminative power and computational efficiency through large-scale simulations across diverse distributional discrepancies (e.g., shifts in location, scale, and higher-order moments) and both two-sample and multi-sample settings. Based on empirical performance, the work proposes a data-characteristic-driven strategy for method selection, establishes a performance ranking, and demonstrates that combining only four to six methods suffices to achieve near-optimal performance in 90%–95% of scenarios.

dataset similaritydistribution comparisonk-sample testing

Human-aligned Quantification of Numerical Data

Nov 14, 2025
AK
Anton Kolonin
🏛️ Novosibirsk State University

This paper addresses the natural quantization of continuous numerical data—automatically identifying statistically meaningful discrete “quantum” intervals whose boundaries align with human intuition. We propose a multi-criteria decision framework integrating the Silhouette coefficient (threshold > 0.65) for assessing inter-class separability, the Dip test (p < 0.5) to validate unimodality assumptions, and an information compression metric to evaluate representational conciseness. Experiments demonstrate that the Silhouette coefficient better captures human perceptual judgments than conventional compression-based approaches, and the synergistic use of all three criteria robustly determines quantization feasibility. We establish, for the first time, interpretable and reproducible quantitative thresholds for quantization validity. A user study confirms strong correlation (r > 0.82) between our metrics and human intuitive judgments. This work provides both theoretical foundations and practical tools for symbolic modeling and interpretable AI.

Assessing metrics for quantifying numerical data into meaningful statesDetermining thresholds for classifying numeric data into distinct categoriesEvaluating correlation between computational metrics and human intuition

How Small Transformation Expose the Weakness of Semantic Similarity Measures

Sep 08, 2025
SL
Serge Lionel Nikiema
🏛️ University of Luxembourg

This study addresses the fundamental question of whether semantic similarity measures genuinely comprehend semantic relationships. We propose the first evaluation framework based on controlled, small-scale semantic transformations to systematically assess the semantic discrimination capability of 18 state-of-the-art methods—including bag-of-words, embedding-based, LLM-based, and structure-aware models—on software engineering texts and code. Experiments reveal that mainstream embedding methods exhibit up to 99.9% misclassification rates in semantic opposition scenarios, exposing their reliance on superficial surface patterns. Substituting cosine similarity for Euclidean distance improves performance by 24–66%. LLM-based methods demonstrate superior fine-grained semantic distinction. Critically, our framework uncovers a foundational limitation in existing measures: their failure to capture semantic essence. It establishes the first reproducible, scalable benchmark paradigm for trustworthy semantic computation in software engineering contexts.

Evaluating semantic similarity measures for software engineering tasksIdentifying flaws where methods confuse opposites and synonymsTesting 18 methods including embeddings and LLMs on semantic understanding

Safire: Similarity Framework for Visualization Retrieval

Oct 18, 2025
HN
Huyen N. Nguyen
🏛️ Harvard Medical School

Current visualisation retrieval lacks a unified conceptual framework for similarity, hindering systematic characterisation along two key dimensions: “what to compare” (similarity criteria) and “how to compare” (representation modalities). This paper introduces Safire, the first two-dimensional theoretical framework that decouples similarity into five criteria—data, visual encoding, interaction, style, and metadata—and four representation modalities: raster images, vector graphics, declarative specifications, and natural language. It further establishes a classification scheme for modalities based on information content and determinacy. Through conceptual modelling, data-centred and human-centred metric analysis, and multi-system case studies, the framework empirically identifies criterion–modality alignment patterns. Safire provides an interpretable, theoretically grounded foundation for multimodal visualisation retrieval, reproducibility, and learning.

Analyzing how representation choices impact visualization retrieval capabilitiesDefining visualization similarity through comparison criteria and representation modalitiesProviding framework recommendations for multimodal learning and AI applications

Hot Scholars

FS

Furao Shen

Department of Computer Science & Technology, Nanjing University
Neural NetworksRobotic Intelligence
SY

Suorong Yang

Nanjing University
Computer VisionDeep LearningMultimodal Learning
KG

Kun Gai

Senior Director & Researcher, Alibaba Group
Machine LearningComputational Advertising
SL

Sheng Liang

CIS LMU Munich & Munich Center for Machine Learning
NLP