Score
Designs and implements resampling-based evaluation protocols and pipelines (e.g., k‑fold, repeated, nested, stratified, grouped, spatio‑temporal, walk‑forward/expanding/rolling-window and constrained variants) that preserve temporal, spatial, class, or group structure to produce realistic out‑of‑sample performance estimates and prevent training–validation leakage. Builds the associated procedures to tune hyperparameters within proper folds, compute and compare evaluation metrics (including turnover/cost adjustments, calibration, predictive coverage and correlation), and run statistical comparisons of models or ensembles.
Model evaluation strategies significantly impact reported accuracy, yet systematic comparisons of cross-validation methods across diverse datasets and algorithms remain limited. Method: This study conducts a comprehensive empirical analysis of hold-out, iterative hold-out, and K-fold cross-validation across two real-world datasets (Framingham Heart Study and COVID-19) and three classifiers (decision tree, naive Bayes, k-nearest neighbors). A parameter sensitivity analysis investigates interactions among test set proportion, random seed, and fold count (K). Contribution/Results: We find that a 10% hold-out split consistently yields higher accuracy than the conventional 20% split across most configurations; iterative hold-out substantially reduces variance; and no universally optimal K exists—its efficacy depends critically on dataset size, feature distribution, and learning algorithm. The study demonstrates that validation strategy selection must jointly account for data characteristics and model properties, providing a reproducible evaluation framework and empirically grounded guidelines for robust model assessment.
This study addresses reproducibility challenges in training histopathology foundation models—arising from software stochasticity, hardware nondeterminism, and inconsistent hyperparameter reporting—by systematically investigating the impact of hyperparameters and data augmentation strategies on model stability. Leveraging the CLIP architecture, we pretrain on QUILT-1M and conduct large-scale ablation studies across three downstream benchmarks: PatchCamelyon, LC25000-Lung, and LC25000-Colon. Key findings include enhanced training consistency with RandomResizedCrop scale ratios of 0.7–0.8, disabling local loss in distributed training, and learning rates below 5.0×10⁻⁵; LC25000-Colon emerges as the most reproducible benchmark. We propose the first reproducibility-oriented best-practice guide specifically for digital pathology modeling, providing methodological foundations for robust development and evaluation of pathology AI systems.
This work addresses the critical challenge of dynamically allocating a limited query budget between resampling and rerouting strategies to maximize the answer accuracy of large language models. Treating these two approaches as competing strategies sharing a common budget, the paper proposes RoR—an online, budget-aware test-time model selection method that dynamically allocates resources based on the marginal gain in accuracy per unit cost. RoR leverages a diverse model pool, online estimation of marginal returns, and a label-agnostic consistency verifier. Evaluated across four heterogeneous benchmarks, it significantly outperforms existing baselines, particularly in settings with high inter-model diversity, and achieves state-of-the-art trade-offs along the cost–accuracy Pareto frontier.
This work addresses the challenge that existing model evaluation methods often fail to reliably assess estimator quality in low-variance settings due to confounding between bias and variance or excessive sensitivity of statistical tests. To overcome this limitation, the authors propose a fault-tolerant evaluation framework that unifies bias and variance modeling through an adjustable tolerance parameter ε, enabling robust assessment of sample-efficient performance estimators within practically acceptable error margins. The framework integrates bias-variance analysis, fault-tolerant evaluation theory, and an adaptive ε-optimization algorithm, making it particularly well-suited for scenarios with low annotation costs. Experimental results demonstrate that the proposed approach provides a more comprehensive and reliable characterization of estimator behavior, significantly enhancing both the practical utility and stability of performance evaluation.
Current continual learning (CL) evaluation protocols suffer from critical flaws—hyperparameter tuning and evaluation are conducted within the same scenario, leading to systematic overestimation of CL capability and employing unrealistic, non-deployable tuning practices. Method: We propose the Generalized Two-stage Evaluation Protocol (GTEP), which strictly decouples hyperparameter optimization (performed solely on a source dataset) from performance evaluation (conducted on a target dataset), thereby enforcing cross-dataset generalization under structurally identical tasks. Contribution/Results: Extensive experiments—over 8,000 runs across CIFAR and ImageNet variants—under both pre-trained and non-pre-trained settings within a class-incremental learning framework demonstrate that mainstream SOTA methods suffer 30–50% average performance degradation under GTEP. This reveals their lack of robustness across deployment scenarios and establishes GTEP as a more rigorous, realistic benchmark for trustworthy continual learning.
Fixed-size benchmarking in model evaluation often fails to balance efficiency, statistical reliability, and diverse objectives, leading to either excessive resource consumption or unreliable results. This work proposes the first adaptive framework that integrates sequential testing into AI model evaluation, dynamically allocating evaluation data based on stopping criteria tailored for model ranking and selection tasks. By combining sequential hypothesis testing, minimum detectable effect analysis, and diminishing returns detection, the method achieves substantial gains in efficiency without compromising rigor. Empirical validation on the Open VLM Leaderboard demonstrates an 80% reduction in computational cost while maintaining a confidence interval width of 2.5 points, significantly enhancing both the practicality and scalability of model evaluation.
Clustering outcomes are highly sensitive to algorithmic choices, preprocessing steps, and the number of clusters, yet conventional validation metrics often fail in high-dimensional, heavy-tailed, or nonlinear biomedical data, leading to irreproducible findings. This work proposes a resampling-driven framework for clustering evaluation that unifies stability and generalization analyses for the first time, enabling diagnostic assessments at global, cluster-level, and sample-level resolutions while producing consensus cluster labels and selection criteria. By circumventing restrictive geometric assumptions inherent in traditional methods, the approach offers a scikit-learn–compatible Python API and a Seurat-compatible R interface. It consistently approximates optimal clustering across six synthetic benchmarks, significantly outperforming existing metrics, and uncovers finer biological structures in real-world genomics and proteomics datasets.