Score
Design and implement evaluation protocols and metric computations for classification and representation models, covering binary and multiclass measures (e.g., precision, recall, AUC), instance‑level and segmentation metrics, re‑identification and regression/classification hybrids, and representation‑distance or spectral metrics. Analyze metric sensitivity and failure modes, compare models across datasets and splits, and produce aggregated performance summaries and diagnostics to guide model selection and improvement.
This study identifies systematic inconsistencies in the implementation of machine learning evaluation metrics across mainstream programming languages—Python, R, and MATLAB—spanning ten task categories: classification, regression, clustering, statistical testing, image segmentation, and image-to-image translation. Through the first large-scale, cross-platform empirical analysis, we quantitatively assess consistency across 100+ metrics. Results reveal that 36 metrics—including Accuracy, AUC, and MAE—are robust across implementations, whereas critical metrics such as Precision, F1-score, IoU, and Within-Cluster Sum of Squares (WCSS) exhibit substantial discrepancies. To address this, we propose the first comprehensive, task-agnostic standardization roadmap for ML evaluation, accompanied by a curated recommendation list. This work provides both theoretical foundations and practical guidelines to enhance cross-platform reproducibility and result reliability in ML research and deployment.
This paper addresses the lack of systematic guidance for selecting loss functions and evaluation metrics in deep learning. Methodologically, it conducts a comprehensive analysis of mathematical properties, gradient behaviors, scale sensitivities, and optimization stability of canonical losses and metrics—including cross-entropy, MSE, IoU, BLEU, F1, and Dice—across 12 mainstream task categories (e.g., regression, classification, CV, NLP). Based on this analysis, it constructs a task-driven “loss–metric alignment matrix.” Its key contribution is the first cross-task, interpretable selection framework that formally bridges the gap between theoretical design principles and empirical engineering practice. The framework provides principled, semantics-aware criteria for matching losses to metrics according to task objectives and optimization dynamics. It has been widely adopted in industry as a standard reference for model development, significantly reducing trial-and-error overhead and improving evaluation consistency across teams and applications.
To address the challenges of imbalanced, heterogeneous multi-source data distributions in real-world scenarios and poor generalization of single-dataset metric learning, this paper proposes Unified Metric Learning (UML)—a novel paradigm for jointly learning a single, robust distance metric across multiple distributions. Methodologically, we introduce PUMA, a parameter-efficient framework that freezes a pretrained backbone, incorporates stochastic adapters and a learnable prompt pool, and integrates contrastive learning with multi-distribution joint optimization to mitigate distributional bias and sample imbalance. Our contributions are threefold: (1) the first UML benchmark comprising eight heterogeneous datasets; (2) a model requiring only 1.4% trainable parameters—69× fewer than state-of-the-art (SOTA) methods—while significantly improving cross-distribution generalization and fairness; and (3) consistent superiority over single-dataset SOTA methods across all tasks on the unified benchmark.
Machine Learning–Enabled Systems (MLES) suffer from poorly quantifiable and unmanageable complexity, hindering systematic governance. Method: This paper proposes the first measurement-driven architectural model for MLES, innovatively extending the ML system reference architecture to natively support automated collection, cross-dimensional correlation analysis, and visualization of multi-faceted complexity indicators—including training data drift, model iteration coupling, and deployment heterogeneity. The approach integrates architectural modeling, software metrics engineering, and complexity quantification theory to shift complexity assessment from qualitative description to quantitative, traceable decision-making. Contribution/Results: Empirical evaluation demonstrates that the model significantly enhances complexity awareness and decision traceability during architectural evolution. It provides both theoretical foundations and practical tooling for sustainable MLES governance, enabling rigorous, evidence-based architectural management throughout the ML lifecycle.
Communication barriers between data scientists and domain experts arise from oversimplified, accuracy-centric model performance reporting, hindering shared understanding of model limitations and contextual applicability. Method: We propose a visualization-mediated model explanation framework grounded in human-computer interaction principles, participatory design, and visual narrative techniques. This yields the first domain-expert-oriented model communication guideline—emphasizing risk, trade-offs, and situational appropriateness rather than isolated metrics like accuracy. An iterative empirical study was conducted using regression models, incorporating structured expert feedback for evaluation. Contribution/Results: The framework significantly improves domain experts’ ability to identify model limitations, recognize inherent trade-offs, and proactively make context-driven adoption decisions. Its core innovation lies in repositioning visualization as an interdisciplinary consensus-building medium—shifting the paradigm from “metric reporting” to “collaborative understanding.”
Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.
This study addresses the challenge posed by severe class imbalance in anomaly detection, which complicates the interpretation and comparison of common evaluation metrics. The authors systematically analyze the behavior of AUROC, AUPR, F1-score, and Matthews Correlation Coefficient (MCC) across varying anomaly ratios and introduce a novel "metric landscape" visualization technique. This approach reveals, for the first time, each metric’s inherent preference for true positive rate versus true negative rate and how their stability varies with imbalance levels. By modeling the relationship between metrics and anomaly prevalence, the work delineates clear applicability boundaries for each metric, thereby providing a principled, interpretable foundation for reliable metric selection in highly imbalanced anomaly detection scenarios.
Existing segmentation evaluation metrics often lack transparency and modularity, making them ill-suited for diverse tasks such as transparent object, specular surface, or lesion segmentation. This work proposes a unified evaluation framework that decomposes metrics into five modular components: prediction representation, target extraction, target matching, score computation, and metric reporting. For the first time, it systematically analyzes the implicit assumptions and design limitations of mainstream binary segmentation metrics through this modular lens. The framework enables task-aware customization of evaluation protocols, reveals evolutionary trajectories among existing metrics, and is accompanied by an open-source toolkit. By offering a principled and interpretable foundation, this approach paves the way for developing more rational and adaptable segmentation evaluation methodologies.
This work addresses the lack of reliable statistical evaluation methods for generative models, which hinders the assessment of their generalization performance and the estimability of evaluation metrics from finite samples. The authors propose a theoretical framework that systematically analyzes the conditions under which common evaluation metrics are statistically estimable, distinguishing between test-class-based metrics and divergence-based metrics in finite-sample settings. Leveraging tools from integral probability metrics (IPMs), Rényi divergences, and fat-shattering dimension, they rigorously establish—for the first time—that IPMs induced by bounded test classes admit arbitrarily accurate estimation from finite samples, whereas KL and Rényi divergences, which depend on rare events, do not. This study provides a foundational theoretical basis and practical guidance for evaluating generative models.
Existing robustness quantification methods rely on generative models and are often constrained by specific architectures or discrete features, limiting their applicability to general discriminative classifiers. This work proposes a novel robustness metric applicable to any probabilistic discriminative classifier and arbitrary feature types, thereby eliminating dependence on generative models and architectural assumptions and offering, for the first time, an effective means to evaluate robustness in generic discriminative models. Building upon this metric, the authors further design a dynamic classifier selection strategy that effectively distinguishes between reliable and unreliable predictions, significantly improving both selection accuracy and broad applicability.