split datasets

Designs and produces dataset partitions for model training, validation, and testing, implementing procedures such as stratified splits to preserve class balance and temporal folds to respect time ordering; ensures appropriate allocation of examples to train, validation, and test sets. Adapts splitting strategies for small or imbalanced datasets (for example by choosing fold sizes, using repeated or nested folds, or controlled resampling) to support reliable model selection and evaluation.

splitdatasets

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.06
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Comparing Cluster-Based Cross-Validation Strategies for Machine Learning Model Evaluation

Jul 29, 2025
AM
Afonso Martini Spezia
🏛️ Universidade Federal do Rio Grande do Sul

Traditional cross-validation often yields biased model performance estimates when data diversity is insufficient. To address this, we propose a novel cross-validation method that integrates Mini-Batch K-Means clustering with class stratification. We systematically evaluate our approach against alternatives—including standard K-Means and hierarchical clustering—across 20 benchmark datasets and four supervised learning models. Results show that the proposed method significantly reduces estimation bias and variance on balanced datasets, outperforming conventional stratified cross-validation; however, traditional stratified CV remains superior on imbalanced data. This work demonstrates the potential of clustering-guided data partitioning to enhance evaluation robustness and provides an interpretable, context-aware framework for selecting appropriate validation strategies under varying data distributions.

Compares bias, variance, and cost on balanced/imbalanced datasetsEvaluates cluster-based cross-validation strategies for model performanceProposes Mini Batch K-Means with class stratification technique

A New Flexible Train-Test Split Algorithm, an approach for choosing among the Hold-out, K-fold cross-validation, and Hold-out iteration

Jan 11, 2025
ZB
Zahra Bami
🏛️ University of Turin | University of Social Welfare and Rehabilitation Science | Macquarie University

Model evaluation strategies significantly impact reported accuracy, yet systematic comparisons of cross-validation methods across diverse datasets and algorithms remain limited. Method: This study conducts a comprehensive empirical analysis of hold-out, iterative hold-out, and K-fold cross-validation across two real-world datasets (Framingham Heart Study and COVID-19) and three classifiers (decision tree, naive Bayes, k-nearest neighbors). A parameter sensitivity analysis investigates interactions among test set proportion, random seed, and fold count (K). Contribution/Results: We find that a 10% hold-out split consistently yields higher accuracy than the conventional 20% split across most configurations; iterative hold-out substantially reduces variance; and no universally optimal K exists—its efficacy depends critically on dataset size, feature distribution, and learning algorithm. The study demonstrates that validation strategy selection must jointly account for data characteristics and model properties, providing a reproducible evaluation framework and empirically grounded guidelines for robust model assessment.

Cross-validation methodsMachine Learning Model EvaluationParameter Configuration

Irredundant k-Fold Cross-Validation

Jul 26, 2025
JS
Jesús S. Aguilar-Ruiz
🏛️ Pablo de Olavide University

Standard k-fold cross-validation suffers from sample reuse: each instance participates in training for $k-1$ folds and testing once, inducing training-set overlap, evaluation bias, and inflated variance estimates. To address this, we propose Single-Use k-Fold Cross-Validation (SU-CV), the first method ensuring each sample is used *exactly once* for training and *exactly once* for testing. SU-CV achieves this via non-overlapping, stratified partitioning—eliminating training-set redundancy while preserving class proportions. It is model-agnostic, requires no architectural modifications, and integrates seamlessly with any classifier. Empirical results demonstrate that SU-CV significantly reduces estimator variance (yielding more conservative performance estimates), mitigates overfitting tendencies, and cuts training computational cost by approximately $(k-1)/k$. Extensive evaluation across multiclass benchmarks confirms its stability and generalization robustness. SU-CV establishes a theoretically sound, fairer, and more efficient benchmark for model evaluation.

Ensures balanced dataset usage to mitigate overfittingProvides consistent performance with lower computational costReduces redundancy in k-fold cross-validation training

This study addresses the distortion in model evaluation caused by conventional random splitting, which often fails to preserve consistent data distributions—such as class imbalance, cluster structure, or spatial autocorrelation—between training and test sets. To mitigate this issue, the authors propose an Optimised-Distribution method that explicitly optimizes the similarity between the distributions of the training and test sets. The approach systematically integrates chi-squared tests, Kolmogorov-Smirnov tests, and Maximum Mean Discrepancy (MMD) as distributional similarity metrics. Evaluated across 15 UCI benchmark datasets against multiple splitting strategies—including random splitting, stratified sampling, Kennard-Stone, Duplex, and SPXY—the proposed method achieves an average MMD similarity of 89.0%, significantly outperforming existing techniques while enhancing both split quality and evaluation stability.

AutoMLdistribution mismatchmodel evaluation

This paper addresses data leakage and computational redundancy arising from preprocessing in partition-based cross-validation. We propose a matrix-algebra-driven acceleration framework that supports model selection for PCA, principal component regression (PCR), and ridge regression. For the first time, we provide provably leakage-free and efficiently verifiable cross-validation algorithms for all 12 essentially distinct combinations of column centering and scaling. Leveraging block-wise preprocessing derivation and validation-set-driven reconstruction of $mathbf{X}^ opmathbf{X}$ and $mathbf{X}^ opmathbf{Y}$, our method reduces both time and space complexity to the order of a single matrix multiplication—scaling independently of the number of folds. An open-source implementation demonstrates that preprocessing overhead remains bounded, numerical precision is preserved, and the framework accommodates all 16 commonly used preprocessing combinations.

Accelerate cross-validation for matrix-based ML modelsReduce redundant computations in training partitionsSupport centering and scaling without data leakage

Latest Papers

What's happening recently
View more

This study investigates whether the value of samples in coreset selection depends on the target learner, specifically examining how decision boundaries between easy-first and geometric coverage criteria shift across models. By conducting frozen-subset intervention and controlled-variable experiments on ImageNet, the authors compare data filtering strategies across varying network widths and architectures, including ResNet and ViT. The findings demonstrate that sample utility exhibits significant learner dependence, with network architecture substantially shifting selection boundaries; notably, coverage criteria consistently dominate under ViT architectures. Furthermore, this work refutes the existence of universal scaling laws for data selection, emphasizing that evaluating filtering strategies necessitates alignment with the target model.

Coreset SelectionFrozen-subset InterventionLearner Dependence

This study addresses the unreliability of conclusions regarding class-imbalance methods derived from single-dataset evaluations. Employing a leakage-free nested cross-validation protocol across 45 binary classification tasks, we conducted large-scale experiments to reassess these techniques. Results reveal that threshold tuning benefits exhibit non-monotonic variation with imbalance ratios and refute the hypothesis that calibration error predicts tuning gains. Furthermore, Random Forest combined with SMOTE demonstrated significant effectiveness across multiple tasks, highlighting the limitations of findings based solely on single fraud datasets. This work systematically clarifies the true utility and applicability boundaries of resampling and threshold tuning, providing robust empirical evidence for the reliable evaluation of imbalanced classification methods.

GeneralizabilityImbalanced ClassificationResampling

Model Class Selection

Nov 14, 2025
RC
Ryan Cecil
🏛️ University of Pittsburgh

This paper addresses the challenge of multi-model-class evaluation by proposing the Model Class Selection (MCS) framework, which identifies the collection of model classes each containing at least one optimal model—thereby enabling formal comparison of performance equivalence across model classes of differing complexity (e.g., interpretable vs. black-box models). MCS generalizes conventional model selection and Model Set Selection (MSS) by integrating likelihood maximization and risk minimization criteria via a data-splitting strategy under mild assumptions. Theoretical analysis establishes its statistical validity. Empirical evaluation—including simulations and real-data experiments—demonstrates that MCS robustly identifies simple, interpretable model classes whose predictive performance matches that of complex models. By bridging interpretability and performance assessment, MCS introduces a novel paradigm and practical tool for explainable AI research.

Compares interpretable models against complex machine learning performanceDevelops data splitting methods for model class selectionGeneralizes model selection to identify optimal model collections

Determining the K in K-fold cross-validation

Nov 16, 2025
KM
Kenichiro McAlinn
🏛️ Temple University | RIKEN Center for Advanced Intelligence Project

Conventional K-fold cross-validation relies on heuristic choices of K (e.g., 5 or 10), leading to suboptimal bias–variance trade-offs in model evaluation. Method: We propose a data- and model-adaptive framework for selecting the optimal K. First, we derive a theoretical upper bound on the estimation uncertainty of cross-validation under finite samples. Then, we formulate a utility-driven optimization objective that explicitly models K-selection as a bias–variance trade-off. Contribution/Results: Empirical validation on real-world datasets—using linear regression and random forests—demonstrates that the optimal K strongly depends on sample size, signal-to-noise ratio, and model complexity; fixed-K conventions thus rest on unwarranted assumptions. Our framework enhances the reliability and interpretability of model evaluation and provides a principled foundation for robust model comparison.

Addressing inability to directly estimate evaluation uncertainty in model validationDetermining optimal K in cross-validation by balancing bias-variance tradeoffReplacing conventional K choices with data-driven principled selection framework

Hot Scholars

XC

Xianhao Chen

Assistant Professor, The University of Hong Kong
Wireless networksmobile edge computingedge AIdistributed learning
SY

Samuel Yen-Chi Chen

Wells Fargo
quantum computationquantum informationmachine learningquantum machine learning
RL

Rayson Laroca

Pontifical Catholic University of Paraná (PUCPR)
Computer VisionDeep LearningPattern Recognition
EH

Elias Hossain

PhD Student, University of Central Florida, USA
(Deep) Machine LearningTrustworthy AILLM ReasoningBioinformatics
NY

Niloofar Yousefi

Assistant Professor
Generative AI for ScienceAI-Guided NanomedicineNext-Gen Therapeutics