Score
Designs and implements algorithms to compute quantitative, distance, ranking and objective metrics from finite or discrete samples, producing automated pipelines for metric extraction and evaluation. Analyzes and builds estimators for these metrics that quantify sampling error, prove robustness to noise or perturbations, and enable discrete scale–space or other finite-sample evaluation procedures.
Evaluation metrics in machine learning are numerous and predominantly relative—dependent on baseline models or data distributions—hindering cross-model and cross-task comparability. Method: We conduct a systematic literature review to comprehensively identify, categorize, and formalize *absolute evaluation metrics*: those with fixed semantic scales, independent of benchmarks or underlying data distributions, for classification, clustering, regression, and ranking tasks. Guided by a task-oriented principle, we construct a structured, unified framework that precisely delineates applicability boundaries and principled selection criteria for each metric. Contribution/Results: This work fills a critical gap in cross-task evaluation guidance, providing practitioners with a standardized, transferable metric selection protocol. By unifying interpretation and usage conventions, it significantly enhances consistency, comparability, and interpretability in model evaluation across diverse learning paradigms.
Traditional model evaluation relies on single-point metrics, failing to characterize performance stability and uncertainty. This paper proposes a small-sample (10–25 runs) uncertainty quantification framework tailored for high-reliability scenarios. It constructs empirical distributions of performance metrics via repeated stochastic experiments—encompassing random data splits, parameter initializations, and hyperparameter perturbations—and robustly estimates confidence intervals for metric quantiles using bias-corrected nonparametric bootstrap combined with quantile regression. To our knowledge, this is the first systematic approach enabling reliable confidence interval estimation for diverse metrics—including accuracy, F1-score, and MAE—in both classification and regression tasks under small-sample regimes. The method achieves high coverage (>90%) while maintaining narrow interval widths, thereby significantly improving robustness in model selection and enhancing decision-making credibility across multiple benchmark datasets.
This work addresses combinatorial bias in classification evaluation metrics under few-shot settings, where group-size disparities induce spurious fairness assessments. Through probabilistic modeling and combinatorial analysis, we systematically identify and quantify the latent bias mechanisms inherent in common metrics—including accuracy and F1-score—under sample-size imbalance. We propose a model-agnostic framework for detecting and correcting such bias, unifying treatment of undefined cases (e.g., zero-denominator scenarios) and metric sensitivity to class distribution. Our approach substantially enhances discriminative power in fairness evaluation under data scarcity, mitigating erroneous attribution and misguided policy interventions arising from metric distortion. The framework provides both theoretical grounding and practical tools for trustworthy AI assessment in resource-constrained environments.
Conventional dimensionality reduction (DR) evaluation suffers from systematic bias due to the frequent adoption of highly correlated metrics, leading to overemphasis on specific structural properties. Method: We propose an empirically grounded metric redundancy reduction framework: first computing Pearson correlation matrices across diverse datasets and DR algorithms; then applying clustering to identify functionally redundant metric groups; and finally retaining only the most representative metric per group—replacing subjective, intent-driven metric selection with objective, behavior-based clustering. Contribution/Results: Our approach significantly improves cross-dataset and cross-algorithm stability of DR evaluations, effectively mitigating structural biases inherent in traditional assessment protocols. Experimental validation demonstrates enhanced reproducibility and generalizability, establishing a principled, data-driven framework for fair and robust comparative evaluation of DR methods.
This study identifies systematic inconsistencies in the implementation of machine learning evaluation metrics across mainstream programming languages—Python, R, and MATLAB—spanning ten task categories: classification, regression, clustering, statistical testing, image segmentation, and image-to-image translation. Through the first large-scale, cross-platform empirical analysis, we quantitatively assess consistency across 100+ metrics. Results reveal that 36 metrics—including Accuracy, AUC, and MAE—are robust across implementations, whereas critical metrics such as Precision, F1-score, IoU, and Within-Cluster Sum of Squares (WCSS) exhibit substantial discrepancies. To address this, we propose the first comprehensive, task-agnostic standardization roadmap for ML evaluation, accompanied by a curated recommendation list. This work provides both theoretical foundations and practical guidelines to enhance cross-platform reproducibility and result reliability in ML research and deployment.
Existing ML evaluation metric libraries suffer from fragmentation, inconsistent implementations, and weak data validation, hindering reliable cross-framework comparisons. To address this, we propose AllMetrics—the first task-aware, standardized metric framework supporting diverse tasks including regression, classification, clustering, segmentation, and image-to-image translation. AllMetrics systematically eliminates implementation discrepancies (IDs) and reporting discrepancies (RDs) via a modular API, strongly typed input validation, multi-task adapters, and cross-language (Python/Matlab/R) consistency verification. Empirical evaluation across healthcare, finance, and real estate domains demonstrates that AllMetrics significantly reduces evaluation errors, improves reproducibility, and enhances trustworthiness in ML workflows.
Real-world distance data are often distorted by noise, missing entries, or violations of the triangle inequality, degrading downstream task performance. This work systematically evaluates the effectiveness of metric repair algorithms, explicitly disentangling two critical subproblems: which edges to repair and how to assign their weights. Through large-scale experiments on both real-world and synthetic non-metric graphs, the study empirically demonstrates—for the first time—that repair efficacy depends primarily on the type and proportion of errors, rather than on error magnitude or graph size. Moreover, accurately identifying the set of edges requiring correction is as crucial as appropriately setting their weights. The findings reveal that most existing repair methods fail to recover a true metric and can even degrade performance, thereby challenging prevailing assumptions in the field.
This work addresses the lack of a universal, tunable, and multi-scenario-compatible metric for data quality assessment, which hinders effective comparison of diverse data cleaning pipelines. To overcome this limitation, the authors propose TOMME—a general-purpose data quality measurement framework based on weighted errors—that extends traditional accuracy into a configurable, composite score. By producing a single quantitative metric, TOMME enables flexible adjustment of error weights according to specific use cases, thereby supporting both automated processing and optimization requirements. Experimental results demonstrate that TOMME exhibits strong adaptability, practicality, and comparability across a variety of scenarios, offering an efficient and unified solution for data quality evaluation and decision-making.