Score
Designs, implements, and validates numerical metrics and procedures that quantify how well classes are separated in a feature or representation space; this covers scale-normalized inter-class distance measures, intra-class variance and spread estimates, margin-based separability scores, and label-free proxies derived from class centers and shapes. Builds dataset-difficulty or representation-quality indices (including RBF-style center/shape metrics) that serve as proxies for downstream linear-probe or logistic-regression performance.
To address the performance degradation of k-nearest neighbors (k-NN) classification on high-dimensional sparse data, this paper proposes a learnable dimension-weighted Minkowski distance metric. By introducing a parameterized dimension-wise weighting mechanism, the method adaptively quantifies each feature’s contribution to similarity computation, thereby dynamically optimizing neighborhood structure. The approach integrates dimension importance analysis with visual interpretability, yielding an end-to-end trainable weighted k-NN framework. Experiments on synthetic datasets and diverse real-world benchmarks—including high-dimensional, small-sample gene expression data—demonstrate that the method significantly improves classification accuracy, achieving an average gain of approximately 10% over standard k-NN. Notably, it exhibits superior robustness in challenging regimes characterized by high dimensionality, severe noise, and limited samples. This work establishes a novel, interpretable, and optimization-friendly distance metric paradigm for high-dimensional pattern classification.
Conventional dimensionality reduction (DR) evaluation suffers from systematic bias due to the frequent adoption of highly correlated metrics, leading to overemphasis on specific structural properties. Method: We propose an empirically grounded metric redundancy reduction framework: first computing Pearson correlation matrices across diverse datasets and DR algorithms; then applying clustering to identify functionally redundant metric groups; and finally retaining only the most representative metric per group—replacing subjective, intent-driven metric selection with objective, behavior-based clustering. Contribution/Results: Our approach significantly improves cross-dataset and cross-algorithm stability of DR evaluations, effectively mitigating structural biases inherent in traditional assessment protocols. Experimental validation demonstrates enhanced reproducibility and generalizability, establishing a principled, data-driven framework for fair and robust comparative evaluation of DR methods.
Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.
Quantifying and interpreting representational similarity between biological systems (e.g., neural activity) and artificial systems (e.g., deep neural networks) remains challenging due to ambiguities in metric choice, interpretability, and functional relevance. Method: We propose the first end-to-end differentiable optimization framework that directly maximizes representational similarity scores between model and neural representations. Using theoretical analysis and synthetic data inversion, we systematically characterize how CKA, angular Procrustes, and normalized Bures similarity (NBS) differentially weight principal component variance. Contributions: We show CKA strongly biases toward high-variance components, whereas angular Procrustes captures low-variance neural dimensions earlier; high similarity scores do not imply functional equivalence, and no universal threshold exists for neuroscientific interpretation; finally, we delineate the feasible score space and hierarchical constraint strengths under multi-metric joint optimization—establishing a theoretical benchmark and practical guidance for representational similarity assessment.
This work addresses the limitations of existing cluster validity indices in unsupervised clustering, where the absence of ground-truth labels and the prevalence of non-convex, irregular, or unevenly distributed data structures hinder reliable evaluation. To overcome these challenges, the authors propose the Central Description Length (CDL) metric, an information-theoretic approach that jointly captures intra-cluster compactness and centroid displacement. CDL estimates an upper bound on the description length of cluster centers, enabling label-free assessment of arbitrary clustering outcomes without reliance on Euclidean distance or kernel functions, thus accommodating clusters of any shape. Empirical results demonstrate that CDL more accurately identifies the true number of clusters and achieves higher Adjusted Rand Index (ARI) scores on synthetic non-convex datasets. Moreover, when applied to embedded representations of MNIST, CIFAR-10, and STL-10, CDL robustly estimates cluster counts closely aligned with the true number of classes.
This work addresses the challenge of selecting high-quality data subsets from noisy labels to achieve performance approaching that of noise-free training. The authors observe that conventional k-nearest neighbors (k-NN) suffer degraded performance in high-dimensional, label-noisy settings and propose a novel approach that integrates symmetry and invariance priors into subset selection. Specifically, they introduce symmetry into the cutstats framework for the first time and theoretically demonstrate that leveraging invariance enables k-NN to asymptotically approach the Bayes optimal classifier. Moreover, they show that even with only partial knowledge of symmetries, effective modeling is achievable through learned symmetry-aware representations. Empirical results confirm that the proposed method substantially improves subset selection quality under high-dimensional label noise, yielding downstream model performance close to that attainable with clean labels.
Existing dimensionality reduction methods often fail to faithfully preserve the relative spatial relationships among class clusters, and conventional evaluation metrics are limited by assumptions of spherical cluster shapes or focus solely on cluster separability, rendering them inadequate for assessing structural fidelity in arbitrarily shaped clusters. To address this, this work proposes CADI, a differentiable metric based on triple-wise inner-angle computation, which introduces an angle-aware mechanism to accurately quantify spatial layout distortions of non-spherical clusters in low-dimensional projections. Building upon CADI, we further develop a dimensionality reduction optimization framework explicitly tailored to preserve cluster structure. Experimental results demonstrate that CADI offers superior interpretability on both real-world and synthetic datasets and significantly enhances the fidelity of cluster organization in visualizations.
In online A/B testing, the relationship between proxy metrics and long-term objectives often breaks down due to user heterogeneity, and relying solely on global correlations can lead to erroneous decisions. This work proposes PROXIMA, a novel framework that introduces a decision-consistency-oriented diagnostic approach for evaluating proxy metrics through three dimensions: normalized effect correlation, directional accuracy, and subgroup vulnerability rate. Integrating causal inference, subgroup analysis, and sensitivity testing, PROXIMA is validated across 80 simulated experiments on the Criteo and KuaiRec datasets. Results show an average decision accuracy of 98.4%; while the subgroup vulnerability rate is markedly higher in recommendation scenarios (68%) than in advertising (13%), directional accuracy exceeds 96% in both, effectively identifying subpopulations where proxy metrics fail.
Existing segmentation evaluation metrics often lack transparency and modularity, making them ill-suited for diverse tasks such as transparent object, specular surface, or lesion segmentation. This work proposes a unified evaluation framework that decomposes metrics into five modular components: prediction representation, target extraction, target matching, score computation, and metric reporting. For the first time, it systematically analyzes the implicit assumptions and design limitations of mainstream binary segmentation metrics through this modular lens. The framework enables task-aware customization of evaluation protocols, reveals evolutionary trajectories among existing metrics, and is accompanied by an open-source toolkit. By offering a principled and interpretable foundation, this approach paves the way for developing more rational and adaptable segmentation evaluation methodologies.