select similarity metric

Designs, implements, and evaluates procedures and models for choosing, computing, or learning similarity and distance measures between data objects (feature vectors and embeddings, sequences, sets, images, strings, and text), including selection rules, metric-learning algorithms, and efficient similarity-computation methods. Analyzes and compares candidate metrics (e.g., cosine, L1/L2, rank-based, and statistical distances), derives selection criteria from data properties (anisotropy, variance concentration, distributional differences), and estimates expected relative performance or improvement when switching metrics.

selectsimilaritymetric

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.35
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$205K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing dataset similarity measures suffer from high computational cost, narrow applicability, sensitivity to data attributes and hyperparameters, and insufficient global robustness. To address these limitations, this paper proposes two novel similarity metrics specifically designed for synthetic data quality assessment and feature selection validation. We introduce the first holistic dataset similarity framework that simultaneously guarantees theoretical soundness, computational efficiency, and parameter robustness. Our approach jointly models probability distances and kernel embeddings by integrating the Maximum Mean Discrepancy (MMD) with geometric consistency constraints—requiring no distributional assumptions and supporting arbitrary-dimensional and heterogeneous data. Evaluated on 12 benchmark datasets, our method achieves an average 37.2% improvement in correlation accuracy over state-of-the-art methods. Moreover, it effectively guides synthetic data generation and feature subset selection.

Computational CostData SimilarityParameter Sensitivity

Methods for quantifying dataset similarity: a review, taxonomy and comparison

Dec 07, 2023
MS
Marieke Stolte
🏛️ TU Dortmund University

This study addresses the critical need for quantifying dataset similarity in model generalization, transfer learning, simulation calibration, and two-sample testing. We systematically survey 118 similarity quantification methods and propose the first ten-dimensional classification framework, organizing approaches into seven technical categories: statistical distances (e.g., Wasserstein distance, Maximum Mean Discrepancy), kernel-based methods, information-theoretic measures, dimensionality-reduction embeddings, permutation tests, generative-model-based discriminators, and Gaussian process likelihood ratios. We develop a multi-dimensional evaluation system balancing theoretical guarantees, interpretability, and practical applicability, yielding a structured recommendation matrix aligned with task requirements and data characteristics. Furthermore, we introduce the first open-source, interactive tool for method selection—enabling real-time filtering and parameter configuration—to significantly enhance both selection efficiency and deployment suitability.

Compare 118 methods on applicability and interpretabilityProvide recommendations for selecting dataset similarity measuresReview and classify methods for quantifying dataset similarity

Conventional dimensionality reduction (DR) evaluation suffers from systematic bias due to the frequent adoption of highly correlated metrics, leading to overemphasis on specific structural properties. Method: We propose an empirically grounded metric redundancy reduction framework: first computing Pearson correlation matrices across diverse datasets and DR algorithms; then applying clustering to identify functionally redundant metric groups; and finally retaining only the most representative metric per group—replacing subjective, intent-driven metric selection with objective, behavior-based clustering. Contribution/Results: Our approach significantly improves cross-dataset and cross-algorithm stability of DR evaluations, effectively mitigating structural biases inherent in traditional assessment protocols. Experimental validation demonstrates enhanced reproducibility and generalizability, establishing a principled, data-driven framework for fair and robust comparative evaluation of DR methods.

Clustering metrics by empirical correlations to avoid overlapImproving stability and reliability of DR projection evaluationsReducing bias in dimensionality reduction evaluation metrics selection

Variance-Adjusted Cosine Distance as Similarity Metric

Feb 04, 2025
SS
Satyajeet Sahoo
🏛️ IIT Kharagpur

Traditional cosine similarity assumes data reside in a Euclidean space and ignores the variance and covariance of random variables, leading to inaccurate similarity estimates when features exhibit correlation and heteroscedasticity. To address this, we propose Variance–Covariance-Corrected Cosine distance (VC-Cosine), the first method to explicitly incorporate second-order statistical structure—i.e., feature variances and covariances—into cosine distance computation. VC-Cosine whitens the feature space via the empirical covariance matrix, thereby adaptively reweighting dimensions according to their correlations and variabilities, and aligning the inner-product geometry with the true underlying data distribution rather than relying on isotropic assumptions. Experiments on the Wisconsin Breast Cancer dataset demonstrate that, when integrated with a k-nearest neighbors classifier, VC-Cosine achieves 100% test accuracy—substantially outperforming standard cosine similarity and other state-of-the-art similarity measures.

Impact of variance and correlation on cosine distance accuracyLimitations of traditional cosine similarity in random variable spaceProposal of a variance-adjusted cosine similarity for improved performance

Learning Unified Distance Metric Across Diverse Data Distributions with Parameter-Efficient Transfer Learning

Sep 16, 2023
SK
Sungyeon Kim
🏛️ Pohang University of Science and Technology (POSTECH) | Korea University

To address the challenges of imbalanced, heterogeneous multi-source data distributions in real-world scenarios and poor generalization of single-dataset metric learning, this paper proposes Unified Metric Learning (UML)—a novel paradigm for jointly learning a single, robust distance metric across multiple distributions. Methodologically, we introduce PUMA, a parameter-efficient framework that freezes a pretrained backbone, incorporates stochastic adapters and a learnable prompt pool, and integrates contrastive learning with multi-distribution joint optimization to mitigate distributional bias and sample imbalance. Our contributions are threefold: (1) the first UML benchmark comprising eight heterogeneous datasets; (2) a model requiring only 1.4% trainable parameters—69× fewer than state-of-the-art (SOTA) methods—while significantly improving cross-distribution generalization and fairness; and (3) consistent superiority over single-dataset SOTA methods across all tasks on the unified benchmark.

Data ImbalanceMetric LearningMulti-Dataset

Latest Papers

What's happening recently
View more

This work proposes a language- and syntax-agnostic approach to string similarity measurement by introducing co-occurrence matrices (COM) and run-length matrices (RLM)—concepts originally from image texture analysis—into the domain of string representation to construct purely statistical, language-independent features. The method integrates multiple statistical measures, including COM, RLM, longest common subsequence, and edit distance. Evaluated on synthetic datasets, COM and RLM significantly outperformed baseline methods in three out of four experiments (p < 0.001). In real-world text plagiarism detection tasks, RLM achieved the best performance, demonstrating the effectiveness and generalizability of the proposed statistical features.

language-independentplagiarism detectionstatistical features

This study addresses the lack of systematic and neutral comparisons among similarity measures for categorical datasets. It presents the first comprehensive evaluation of several prominent methods—including edge-count tests, constrained minimum distance, graph-based tests, Classifier Two-Sample Tests (C2ST), and the Maximum Mean Discrepancy with Categorical Metrics (MMCM)—assessing their ability to detect distributional differences and their computational costs in both two-sample and multi-sample settings. The results demonstrate that the Friedman–Rafsky test achieves the best overall performance in two-sample tasks, while MMCM excels in multi-sample scenarios by offering both high statistical power and computational efficiency. This work provides empirical evidence and practical guidance for selecting appropriate similarity measures when analyzing categorical data.

categorical datadataset similaritymulti-sample comparison

This study addresses the lack of systematic and impartial evaluation of existing methods for measuring distributional similarity in numerical data, which hinders informed selection in practice. The authors construct the first comprehensive benchmarking framework encompassing 36 similarity measures for continuous data—including statistical tests, distance-based metrics, and embedding approaches—and evaluate their discriminative power and computational efficiency through large-scale simulations across diverse distributional discrepancies (e.g., shifts in location, scale, and higher-order moments) and both two-sample and multi-sample settings. Based on empirical performance, the work proposes a data-characteristic-driven strategy for method selection, establishes a performance ranking, and demonstrates that combining only four to six methods suffices to achieve near-optimal performance in 90%–95% of scenarios.

dataset similaritydistribution comparisonk-sample testing

Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.

class imbalancemodel evaluationperformance metrics

This study addresses the limitations of evaluating multiclass classifiers using single performance metrics, which often leads to misleading conclusions. To overcome this, the work proposes a multidimensional evaluation paradigm that leverages the PyCM library to construct a comprehensive analytical framework, enabling systematic comparison of classifier performance across a diverse set of evaluation metrics. Through two case studies, the research uncovers nuanced performance trade-offs that conventional metrics fail to capture, thereby demonstrating the necessity and effectiveness of multidimensional assessment in model selection and optimization. The findings further highlight the unique value of PyCM in facilitating thorough and precise evaluation of multiclass classification systems.

classifier comparisonevaluation frameworkmodel evaluation

Hot Scholars

TE

Tome Eftimov

Computer Systems Department, Jožef Stefan Institute
StatisticsStochastic Optimization AlgorithmsMachine learningNatural Language Processing
DZ

Dongzhan Zhou

Researcher at Shanghai AI Lab
AI4Sciencecomputer visiondeep learning
KC

Kai Chen

Shanghai AI Laboratory
LLMVLMComputer Vision
LZ

Linfeng Zhang

DP Technology; AI for Science Institute
AI for Sciencemulti-scale modelingmolecular simulationdrug/materials design
YZ

Yuhang Zang

Shanghai AI Laboratory
Natural Language ProcessingVision Language Model