Comparing Cluster-Based Cross-Validation Strategies for Machine Learning Model Evaluation

📅 2025-07-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Traditional cross-validation often yields biased model performance estimates when data diversity is insufficient. To address this, we propose a novel cross-validation method that integrates Mini-Batch K-Means clustering with class stratification. We systematically evaluate our approach against alternatives—including standard K-Means and hierarchical clustering—across 20 benchmark datasets and four supervised learning models. Results show that the proposed method significantly reduces estimation bias and variance on balanced datasets, outperforming conventional stratified cross-validation; however, traditional stratified CV remains superior on imbalanced data. This work demonstrates the potential of clustering-guided data partitioning to enhance evaluation robustness and provides an interpretable, context-aware framework for selecting appropriate validation strategies under varying data distributions.

Technology Category

Machine Learning: Ensemble MethodsData Mining & Knowledge Management: Anomaly/Outlier DetectionComputer Vision: Segmentation

Application Category

Web Mining and Content Analysis: Robustness and generalizability of Web mining methodsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
Cross-validation plays a fundamental role in Machine Learning, enabling robust evaluation of model performance and preventing overestimation on training and validation data. However, one of its drawbacks is the potential to create data subsets (folds) that do not adequately represent the diversity of the original dataset, which can lead to biased performance estimates. The objective of this work is to deepen the investigation of cluster-based cross-validation strategies by analyzing the performance of different clustering algorithms through experimental comparison. Additionally, a new cross-validation technique that combines Mini Batch K-Means with class stratification is proposed. Experiments were conducted on 20 datasets (both balanced and imbalanced) using four supervised learning algorithms, comparing cross-validation strategies in terms of bias, variance, and computational cost. The technique that uses Mini Batch K-Means with class stratification outperformed others in terms of bias and variance on balanced datasets, though it did not significantly reduce computational cost. On imbalanced datasets, traditional stratified cross-validation consistently performed better, showing lower bias, variance, and computational cost, making it a safe choice for performance evaluation in scenarios with class imbalance. In the comparison of different clustering algorithms, no single algorithm consistently stood out as superior. Overall, this work contributes to improving predictive model evaluation strategies by providing a deeper understanding of the potential of cluster-based data splitting techniques and reaffirming the effectiveness of well-established strategies like stratified cross-validation. Moreover, it highlights perspectives for increasing the robustness and reliability of model evaluations, especially in datasets with clustering characteristics.
Problem

Research questions and friction points this paper is trying to address.

Evaluates cluster-based cross-validation strategies for model performance
Proposes Mini Batch K-Means with class stratification technique
Compares bias, variance, and cost on balanced/imbalanced datasets
Innovation

Methods, ideas, or system contributions that make the work stand out.

Proposes Mini Batch K-Means with class stratification
Compares clustering algorithms for cross-validation
Evaluates bias, variance, and computational cost
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
A
Afonso Martini Spezia
Instituto de Informática, Universidade Federal do Rio Grande do Sul (UFRGS), Porto Alegre, Brazil
Mariana Recamonde-Mendoza
Mariana Recamonde-Mendoza
Universidade Federal do Rio Grande do Sul/Hospital de Clínicas de Porto Alegre
Machine LearningData ScienceBioinformaticsComputational Biology