cross-validation

Designs and implements resampling-based evaluation protocols and pipelines (e.g., k‑fold, repeated, nested, stratified, grouped, spatio‑temporal, walk‑forward/expanding/rolling-window and constrained variants) that preserve temporal, spatial, class, or group structure to produce realistic out‑of‑sample performance estimates and prevent training–validation leakage. Builds the associated procedures to tune hyperparameters within proper folds, compute and compare evaluation metrics (including turnover/cost adjustments, calibration, predictive coverage and correlation), and run statistical comparisons of models or ensembles.

cross-validation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.57
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$207K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A New Flexible Train-Test Split Algorithm, an approach for choosing among the Hold-out, K-fold cross-validation, and Hold-out iteration

Jan 11, 2025
ZB
Zahra Bami
🏛️ University of Turin | University of Social Welfare and Rehabilitation Science | Macquarie University

Model evaluation strategies significantly impact reported accuracy, yet systematic comparisons of cross-validation methods across diverse datasets and algorithms remain limited. Method: This study conducts a comprehensive empirical analysis of hold-out, iterative hold-out, and K-fold cross-validation across two real-world datasets (Framingham Heart Study and COVID-19) and three classifiers (decision tree, naive Bayes, k-nearest neighbors). A parameter sensitivity analysis investigates interactions among test set proportion, random seed, and fold count (K). Contribution/Results: We find that a 10% hold-out split consistently yields higher accuracy than the conventional 20% split across most configurations; iterative hold-out substantially reduces variance; and no universally optimal K exists—its efficacy depends critically on dataset size, feature distribution, and learning algorithm. The study demonstrates that validation strategy selection must jointly account for data characteristics and model properties, providing a reproducible evaluation framework and empirically grounded guidelines for robust model assessment.

Cross-validation methodsMachine Learning Model EvaluationParameter Configuration

Hyperparameter Optimization and Reproducibility in Deep Learning Model Training

Oct 16, 2025
UA
Usman Afzaal
🏛️ The Ohio State University Wexner Medical Center | Wake Forest University School of Medicine | Wake Forest Institute for Regenerative Medicine

This study addresses reproducibility challenges in training histopathology foundation models—arising from software stochasticity, hardware nondeterminism, and inconsistent hyperparameter reporting—by systematically investigating the impact of hyperparameters and data augmentation strategies on model stability. Leveraging the CLIP architecture, we pretrain on QUILT-1M and conduct large-scale ablation studies across three downstream benchmarks: PatchCamelyon, LC25000-Lung, and LC25000-Colon. Key findings include enhanced training consistency with RandomResizedCrop scale ratios of 0.7–0.8, disabling local loss in distributed training, and learning rates below 5.0×10⁻⁵; LC25000-Colon emerges as the most reproducible benchmark. We propose the first reproducibility-oriented best-practice guide specifically for digital pathology modeling, providing methodological foundations for robust development and evaluation of pathology AI systems.

Addressing reproducibility challenges in histopathology foundation model trainingEvaluating hyperparameter impacts on model performance across datasetsIdentifying optimal configurations for stable computational pathology models

This work addresses the critical challenge of dynamically allocating a limited query budget between resampling and rerouting strategies to maximize the answer accuracy of large language models. Treating these two approaches as competing strategies sharing a common budget, the paper proposes RoR—an online, budget-aware test-time model selection method that dynamically allocates resources based on the marginal gain in accuracy per unit cost. RoR leverages a diverse model pool, online estimation of marginal returns, and a label-agnostic consistency verifier. Evaluated across four heterogeneous benchmarks, it significantly outperforms existing baselines, particularly in settings with high inter-model diversity, and achieves state-of-the-art trade-offs along the cost–accuracy Pareto frontier.

budget-awarelarge language modelsrerouting

This work addresses the challenge that existing model evaluation methods often fail to reliably assess estimator quality in low-variance settings due to confounding between bias and variance or excessive sensitivity of statistical tests. To overcome this limitation, the authors propose a fault-tolerant evaluation framework that unifies bias and variance modeling through an adjustable tolerance parameter ε, enabling robust assessment of sample-efficient performance estimators within practically acceptable error margins. The framework integrates bias-variance analysis, fault-tolerant evaluation theory, and an adaptive ε-optimization algorithm, making it particularly well-suited for scenarios with low annotation costs. Experimental results demonstrate that the proposed approach provides a more comprehensive and reliable characterization of estimator behavior, significantly enhancing both the practical utility and stability of performance evaluation.

bias-variance tradeofffault-tolerant evaluationmodel performance estimation

Hyperparameters in Continual Learning: a Reality Check

Mar 14, 2024
SC
Sungmin Cha
🏛️ New York University | Genentech

Current continual learning (CL) evaluation protocols suffer from critical flaws—hyperparameter tuning and evaluation are conducted within the same scenario, leading to systematic overestimation of CL capability and employing unrealistic, non-deployable tuning practices. Method: We propose the Generalized Two-stage Evaluation Protocol (GTEP), which strictly decouples hyperparameter optimization (performed solely on a source dataset) from performance evaluation (conducted on a target dataset), thereby enforcing cross-dataset generalization under structurally identical tasks. Contribution/Results: Extensive experiments—over 8,000 runs across CIFAR and ImageNet variants—under both pre-trained and non-pre-trained settings within a class-incremental learning framework demonstrate that mainstream SOTA methods suffer 30–50% average performance degradation under GTEP. This reveals their lack of robustness across deployment scenarios and establishes GTEP as a more rigorous, realistic benchmark for trustworthy continual learning.

Challenges unrealistic tuning in conventional continual learning protocolsEvaluates hyperparameter generalizability in continual learning scenariosProposes GTEP to assess algorithm performance across unseen datasets

Latest Papers

What's happening recently
View more

Fixed-size benchmarking in model evaluation often fails to balance efficiency, statistical reliability, and diverse objectives, leading to either excessive resource consumption or unreliable results. This work proposes the first adaptive framework that integrates sequential testing into AI model evaluation, dynamically allocating evaluation data based on stopping criteria tailored for model ranking and selection tasks. By combining sequential hypothesis testing, minimum detectable effect analysis, and diminishing returns detection, the method achieves substantial gains in efficiency without compromising rigor. Empirical validation on the Open VLM Leaderboard demonstrates an 80% reduction in computational cost while maintaining a confidence interval width of 2.5 points, significantly enhancing both the practicality and scalability of model evaluation.

computational costfixed-size benchmarksmodel evaluation

Clustering outcomes are highly sensitive to algorithmic choices, preprocessing steps, and the number of clusters, yet conventional validation metrics often fail in high-dimensional, heavy-tailed, or nonlinear biomedical data, leading to irreproducible findings. This work proposes a resampling-driven framework for clustering evaluation that unifies stability and generalization analyses for the first time, enabling diagnostic assessments at global, cluster-level, and sample-level resolutions while producing consensus cluster labels and selection criteria. By circumventing restrictive geometric assumptions inherent in traditional methods, the approach offers a scikit-learn–compatible Python API and a Seurat-compatible R interface. It consistently approximates optimal clustering across six synthetic benchmarks, significantly outperforming existing metrics, and uncovers finer biological structures in real-world genomics and proteomics datasets.

cluster stabilityclustering validationhigh-dimensional data

Hot Scholars

UB

Ulas Bagci

Northwestern University
artificial intelligencedeep learningbiomedical image analysismedical image computing
JL

Jing Lei

Carnegie Mellon University
Probability and Statistics
FF

Faika Fairuj Preotee

Adjunct Lecturer of Department of CSE, Southeast University, Bangladesh
Computer VisionNatural Language Processing
SR

Shamim Rahim Refat

Software Engineer Intern (AI/ML), Technohaven Company Ltd.
Computer VisionNatural Language Processing
SS

Shuvashis Sarker

Adjunct Lecturer of Department of CSE, Southeast University, Bangladesh
Computer VisionNatural Language ProcessingExplainable AiSignal Processing