Score
Designs and produces dataset partitions for model training, validation, and testing, implementing procedures such as stratified splits to preserve class balance and temporal folds to respect time ordering; ensures appropriate allocation of examples to train, validation, and test sets. Adapts splitting strategies for small or imbalanced datasets (for example by choosing fold sizes, using repeated or nested folds, or controlled resampling) to support reliable model selection and evaluation.
Traditional cross-validation often yields biased model performance estimates when data diversity is insufficient. To address this, we propose a novel cross-validation method that integrates Mini-Batch K-Means clustering with class stratification. We systematically evaluate our approach against alternatives—including standard K-Means and hierarchical clustering—across 20 benchmark datasets and four supervised learning models. Results show that the proposed method significantly reduces estimation bias and variance on balanced datasets, outperforming conventional stratified cross-validation; however, traditional stratified CV remains superior on imbalanced data. This work demonstrates the potential of clustering-guided data partitioning to enhance evaluation robustness and provides an interpretable, context-aware framework for selecting appropriate validation strategies under varying data distributions.
Model evaluation strategies significantly impact reported accuracy, yet systematic comparisons of cross-validation methods across diverse datasets and algorithms remain limited. Method: This study conducts a comprehensive empirical analysis of hold-out, iterative hold-out, and K-fold cross-validation across two real-world datasets (Framingham Heart Study and COVID-19) and three classifiers (decision tree, naive Bayes, k-nearest neighbors). A parameter sensitivity analysis investigates interactions among test set proportion, random seed, and fold count (K). Contribution/Results: We find that a 10% hold-out split consistently yields higher accuracy than the conventional 20% split across most configurations; iterative hold-out substantially reduces variance; and no universally optimal K exists—its efficacy depends critically on dataset size, feature distribution, and learning algorithm. The study demonstrates that validation strategy selection must jointly account for data characteristics and model properties, providing a reproducible evaluation framework and empirically grounded guidelines for robust model assessment.
Standard k-fold cross-validation suffers from sample reuse: each instance participates in training for $k-1$ folds and testing once, inducing training-set overlap, evaluation bias, and inflated variance estimates. To address this, we propose Single-Use k-Fold Cross-Validation (SU-CV), the first method ensuring each sample is used *exactly once* for training and *exactly once* for testing. SU-CV achieves this via non-overlapping, stratified partitioning—eliminating training-set redundancy while preserving class proportions. It is model-agnostic, requires no architectural modifications, and integrates seamlessly with any classifier. Empirical results demonstrate that SU-CV significantly reduces estimator variance (yielding more conservative performance estimates), mitigates overfitting tendencies, and cuts training computational cost by approximately $(k-1)/k$. Extensive evaluation across multiclass benchmarks confirms its stability and generalization robustness. SU-CV establishes a theoretically sound, fairer, and more efficient benchmark for model evaluation.
This study addresses the distortion in model evaluation caused by conventional random splitting, which often fails to preserve consistent data distributions—such as class imbalance, cluster structure, or spatial autocorrelation—between training and test sets. To mitigate this issue, the authors propose an Optimised-Distribution method that explicitly optimizes the similarity between the distributions of the training and test sets. The approach systematically integrates chi-squared tests, Kolmogorov-Smirnov tests, and Maximum Mean Discrepancy (MMD) as distributional similarity metrics. Evaluated across 15 UCI benchmark datasets against multiple splitting strategies—including random splitting, stratified sampling, Kennard-Stone, Duplex, and SPXY—the proposed method achieves an average MMD similarity of 89.0%, significantly outperforming existing techniques while enhancing both split quality and evaluation stability.
This paper addresses data leakage and computational redundancy arising from preprocessing in partition-based cross-validation. We propose a matrix-algebra-driven acceleration framework that supports model selection for PCA, principal component regression (PCR), and ridge regression. For the first time, we provide provably leakage-free and efficiently verifiable cross-validation algorithms for all 12 essentially distinct combinations of column centering and scaling. Leveraging block-wise preprocessing derivation and validation-set-driven reconstruction of $mathbf{X}^ opmathbf{X}$ and $mathbf{X}^ opmathbf{Y}$, our method reduces both time and space complexity to the order of a single matrix multiplication—scaling independently of the number of folds. An open-source implementation demonstrates that preprocessing overhead remains bounded, numerical precision is preserved, and the framework accommodates all 16 commonly used preprocessing combinations.
论文提出使用预测碎片化方法来解决测试时适应性问题,通过测量初始模型与适应后模型之间的不一致性,以减少有害接受区域,提高适应效果。
This study investigates whether the value of samples in coreset selection depends on the target learner, specifically examining how decision boundaries between easy-first and geometric coverage criteria shift across models. By conducting frozen-subset intervention and controlled-variable experiments on ImageNet, the authors compare data filtering strategies across varying network widths and architectures, including ResNet and ViT. The findings demonstrate that sample utility exhibits significant learner dependence, with network architecture substantially shifting selection boundaries; notably, coverage criteria consistently dominate under ViT architectures. Furthermore, this work refutes the existence of universal scaling laws for data selection, emphasizing that evaluating filtering strategies necessitates alignment with the target model.
This study addresses the unreliability of conclusions regarding class-imbalance methods derived from single-dataset evaluations. Employing a leakage-free nested cross-validation protocol across 45 binary classification tasks, we conducted large-scale experiments to reassess these techniques. Results reveal that threshold tuning benefits exhibit non-monotonic variation with imbalance ratios and refute the hypothesis that calibration error predicts tuning gains. Furthermore, Random Forest combined with SMOTE demonstrated significant effectiveness across multiple tasks, highlighting the limitations of findings based solely on single fraud datasets. This work systematically clarifies the true utility and applicability boundaries of resampling and threshold tuning, providing robust empirical evidence for the reliable evaluation of imbalanced classification methods.
This paper addresses the challenge of multi-model-class evaluation by proposing the Model Class Selection (MCS) framework, which identifies the collection of model classes each containing at least one optimal model—thereby enabling formal comparison of performance equivalence across model classes of differing complexity (e.g., interpretable vs. black-box models). MCS generalizes conventional model selection and Model Set Selection (MSS) by integrating likelihood maximization and risk minimization criteria via a data-splitting strategy under mild assumptions. Theoretical analysis establishes its statistical validity. Empirical evaluation—including simulations and real-data experiments—demonstrates that MCS robustly identifies simple, interpretable model classes whose predictive performance matches that of complex models. By bridging interpretability and performance assessment, MCS introduces a novel paradigm and practical tool for explainable AI research.
Conventional K-fold cross-validation relies on heuristic choices of K (e.g., 5 or 10), leading to suboptimal bias–variance trade-offs in model evaluation. Method: We propose a data- and model-adaptive framework for selecting the optimal K. First, we derive a theoretical upper bound on the estimation uncertainty of cross-validation under finite samples. Then, we formulate a utility-driven optimization objective that explicitly models K-selection as a bias–variance trade-off. Contribution/Results: Empirical validation on real-world datasets—using linear regression and random forests—demonstrates that the optimal K strongly depends on sample size, signal-to-noise ratio, and model complexity; fixed-K conventions thus rest on unwarranted assumptions. Our framework enhances the reliability and interpretability of model evaluation and provides a principled foundation for robust model comparison.