Score
Designs, implements, and evaluates procedures and models for choosing, computing, or learning similarity and distance measures between data objects (feature vectors and embeddings, sequences, sets, images, strings, and text), including selection rules, metric-learning algorithms, and efficient similarity-computation methods. Analyzes and compares candidate metrics (e.g., cosine, L1/L2, rank-based, and statistical distances), derives selection criteria from data properties (anisotropy, variance concentration, distributional differences), and estimates expected relative performance or improvement when switching metrics.
Existing dataset similarity measures suffer from high computational cost, narrow applicability, sensitivity to data attributes and hyperparameters, and insufficient global robustness. To address these limitations, this paper proposes two novel similarity metrics specifically designed for synthetic data quality assessment and feature selection validation. We introduce the first holistic dataset similarity framework that simultaneously guarantees theoretical soundness, computational efficiency, and parameter robustness. Our approach jointly models probability distances and kernel embeddings by integrating the Maximum Mean Discrepancy (MMD) with geometric consistency constraints—requiring no distributional assumptions and supporting arbitrary-dimensional and heterogeneous data. Evaluated on 12 benchmark datasets, our method achieves an average 37.2% improvement in correlation accuracy over state-of-the-art methods. Moreover, it effectively guides synthetic data generation and feature subset selection.
This study addresses the critical need for quantifying dataset similarity in model generalization, transfer learning, simulation calibration, and two-sample testing. We systematically survey 118 similarity quantification methods and propose the first ten-dimensional classification framework, organizing approaches into seven technical categories: statistical distances (e.g., Wasserstein distance, Maximum Mean Discrepancy), kernel-based methods, information-theoretic measures, dimensionality-reduction embeddings, permutation tests, generative-model-based discriminators, and Gaussian process likelihood ratios. We develop a multi-dimensional evaluation system balancing theoretical guarantees, interpretability, and practical applicability, yielding a structured recommendation matrix aligned with task requirements and data characteristics. Furthermore, we introduce the first open-source, interactive tool for method selection—enabling real-time filtering and parameter configuration—to significantly enhance both selection efficiency and deployment suitability.
Conventional dimensionality reduction (DR) evaluation suffers from systematic bias due to the frequent adoption of highly correlated metrics, leading to overemphasis on specific structural properties. Method: We propose an empirically grounded metric redundancy reduction framework: first computing Pearson correlation matrices across diverse datasets and DR algorithms; then applying clustering to identify functionally redundant metric groups; and finally retaining only the most representative metric per group—replacing subjective, intent-driven metric selection with objective, behavior-based clustering. Contribution/Results: Our approach significantly improves cross-dataset and cross-algorithm stability of DR evaluations, effectively mitigating structural biases inherent in traditional assessment protocols. Experimental validation demonstrates enhanced reproducibility and generalizability, establishing a principled, data-driven framework for fair and robust comparative evaluation of DR methods.
Traditional cosine similarity assumes data reside in a Euclidean space and ignores the variance and covariance of random variables, leading to inaccurate similarity estimates when features exhibit correlation and heteroscedasticity. To address this, we propose Variance–Covariance-Corrected Cosine distance (VC-Cosine), the first method to explicitly incorporate second-order statistical structure—i.e., feature variances and covariances—into cosine distance computation. VC-Cosine whitens the feature space via the empirical covariance matrix, thereby adaptively reweighting dimensions according to their correlations and variabilities, and aligning the inner-product geometry with the true underlying data distribution rather than relying on isotropic assumptions. Experiments on the Wisconsin Breast Cancer dataset demonstrate that, when integrated with a k-nearest neighbors classifier, VC-Cosine achieves 100% test accuracy—substantially outperforming standard cosine similarity and other state-of-the-art similarity measures.
To address the challenges of imbalanced, heterogeneous multi-source data distributions in real-world scenarios and poor generalization of single-dataset metric learning, this paper proposes Unified Metric Learning (UML)—a novel paradigm for jointly learning a single, robust distance metric across multiple distributions. Methodologically, we introduce PUMA, a parameter-efficient framework that freezes a pretrained backbone, incorporates stochastic adapters and a learnable prompt pool, and integrates contrastive learning with multi-distribution joint optimization to mitigate distributional bias and sample imbalance. Our contributions are threefold: (1) the first UML benchmark comprising eight heterogeneous datasets; (2) a model requiring only 1.4% trainable parameters—69× fewer than state-of-the-art (SOTA) methods—while significantly improving cross-distribution generalization and fairness; and (3) consistent superiority over single-dataset SOTA methods across all tasks on the unified benchmark.
This work proposes a language- and syntax-agnostic approach to string similarity measurement by introducing co-occurrence matrices (COM) and run-length matrices (RLM)—concepts originally from image texture analysis—into the domain of string representation to construct purely statistical, language-independent features. The method integrates multiple statistical measures, including COM, RLM, longest common subsequence, and edit distance. Evaluated on synthetic datasets, COM and RLM significantly outperformed baseline methods in three out of four experiments (p < 0.001). In real-world text plagiarism detection tasks, RLM achieved the best performance, demonstrating the effectiveness and generalizability of the proposed statistical features.
This study addresses the lack of systematic and neutral comparisons among similarity measures for categorical datasets. It presents the first comprehensive evaluation of several prominent methods—including edge-count tests, constrained minimum distance, graph-based tests, Classifier Two-Sample Tests (C2ST), and the Maximum Mean Discrepancy with Categorical Metrics (MMCM)—assessing their ability to detect distributional differences and their computational costs in both two-sample and multi-sample settings. The results demonstrate that the Friedman–Rafsky test achieves the best overall performance in two-sample tasks, while MMCM excels in multi-sample scenarios by offering both high statistical power and computational efficiency. This work provides empirical evidence and practical guidance for selecting appropriate similarity measures when analyzing categorical data.
This study addresses the lack of systematic and impartial evaluation of existing methods for measuring distributional similarity in numerical data, which hinders informed selection in practice. The authors construct the first comprehensive benchmarking framework encompassing 36 similarity measures for continuous data—including statistical tests, distance-based metrics, and embedding approaches—and evaluate their discriminative power and computational efficiency through large-scale simulations across diverse distributional discrepancies (e.g., shifts in location, scale, and higher-order moments) and both two-sample and multi-sample settings. Based on empirical performance, the work proposes a data-characteristic-driven strategy for method selection, establishes a performance ranking, and demonstrates that combining only four to six methods suffices to achieve near-optimal performance in 90%–95% of scenarios.
Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.
This study addresses the limitations of evaluating multiclass classifiers using single performance metrics, which often leads to misleading conclusions. To overcome this, the work proposes a multidimensional evaluation paradigm that leverages the PyCM library to construct a comprehensive analytical framework, enabling systematic comparison of classifier performance across a diverse set of evaluation metrics. Through two case studies, the research uncovers nuanced performance trade-offs that conventional metrics fail to capture, thereby demonstrating the necessity and effectiveness of multidimensional assessment in model selection and optimization. The findings further highlight the unique value of PyCM in facilitating thorough and precise evaluation of multiclass classification systems.