Machine Learning Evaluation Metric Discrepancies across Programming Languages and Their Components: Need for Standardization

📅 2024-11-18
🏛️ arXiv.org
📈 Citations: 0
Influential: 0
📄 PDF

career value

172K/year
🤖 AI Summary
This study identifies systematic inconsistencies in the implementation of machine learning evaluation metrics across mainstream programming languages—Python, R, and MATLAB—spanning ten task categories: classification, regression, clustering, statistical testing, image segmentation, and image-to-image translation. Through the first large-scale, cross-platform empirical analysis, we quantitatively assess consistency across 100+ metrics. Results reveal that 36 metrics—including Accuracy, AUC, and MAE—are robust across implementations, whereas critical metrics such as Precision, F1-score, IoU, and Within-Cluster Sum of Squares (WCSS) exhibit substantial discrepancies. To address this, we propose the first comprehensive, task-agnostic standardization roadmap for ML evaluation, accompanied by a curated recommendation list. This work provides both theoretical foundations and practical guidelines to enhance cross-platform reproducibility and result reliability in ML research and deployment.

Technology Category

Application Category

📝 Abstract
This study evaluates metrics for tasks such as classification, regression, clustering, correlation analysis, statistical tests, segmentation, and image-to-image (I2I) translation. Metrics were compared across Python libraries, R packages, and Matlab functions to assess their consistency and highlight discrepancies. The findings underscore the need for a unified roadmap to standardize metrics, ensuring reliable and reproducible ML evaluations across platforms. This study examined a wide range of evaluation metrics across various tasks and found only some to be consistent across platforms, such as (i) Accuracy, Balanced Accuracy, Cohens Kappa, F-beta Score, MCC, Geometric Mean, AUC, and Log Loss in binary classification; (ii) Accuracy, Cohens Kappa, and F-beta Score in multi-class classification; (iii) MAE, MSE, RMSE, MAPE, Explained Variance, Median AE, MSLE, and Huber in regression; (iv) Davies-Bouldin Index and Calinski-Harabasz Index in clustering; (v) Pearson, Spearman, Kendall's Tau, Mutual Information, Distance Correlation, Percbend, Shepherd, and Partial Correlation in correlation analysis; (vi) Paired t-test, Chi-Square Test, ANOVA, Kruskal-Wallis Test, Shapiro-Wilk Test, Welchs t-test, and Bartlett's test in statistical tests; (vii) Accuracy, Precision, and Recall in 2D segmentation; (viii) Accuracy in 3D segmentation; (ix) MAE, MSE, RMSE, and R-Squared in 2D-I2I translation; and (x) MAE, MSE, and RMSE in 3D-I2I translation. Given observation of discrepancies in a number of metrics (e.g. precision, recall and F1 score in binary classification, WCSS in clustering, multiple statistical tests, and IoU in segmentation, amongst multiple metrics), this study concludes that ML evaluation metrics require standardization and recommends that future research use consistent metrics for different tasks to effectively compare ML techniques and solutions.
Problem

Research questions and friction points this paper is trying to address.

Evaluates discrepancies in ML metrics across Python, R, and Matlab.
Highlights inconsistencies in metrics for classification, regression, and clustering.
Advocates for standardization to ensure reliable ML evaluations.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Standardize ML metrics across programming languages
Compare metrics in Python, R, Matlab for consistency
Highlight discrepancies in ML evaluation metrics
🔎 Similar Papers
No similar papers found.
M
Mohammad R. Salmanpour
Department of Radiology, University of British Columbia, Vancouver BC, Canada
M
Morteza Alizadeh
Department of Mathematics, University of Isfahan, Isfahan, Iran
G
Ghazal Mousavi
School of Electrical and Computer Engineering, University of Tehran, Tehran, Iran
S
Saba Sadeghi
Department of Statistics, Shiraz University, Shiraz, Iran
S
Sajad Amiri
Technological Virtual Collaboration (TECVICO CORP.), Vancouver, BC, Canada
M
M. Oveisi
Department of Computer Science, University of British Columbia, Vancouver, BC, Canada
Arman Rahmim
Arman Rahmim
Professor of Radiology, Physics and Biomedical Engineering, University of British Columbia
computational imagingmolecular imagingpersonalized cancer therapyAItheranostics
Ilker Hacihaliloglu
Ilker Hacihaliloglu
Department of Radiology, Department of Medicine, University of British Columbia
Biomedical EngineeringMedical Image ProcessingUltrasound Image ProcessingImage Guided Surgery and TherapyDeep Learning f