Score
Designs and implements cross-validation and holdout schemes that treat entire measurement replicates as the unit of validation, holding out all samples from one replicate as test data (leave-one-replicate-out / replicate-aware holdout). Uses nested leave-one-replicate-out procedures when tuning or selecting models to prevent replicate-wise information leakage and to produce realistic estimates of generalization across replicates.
Cross-validation data splitting in high-dimensional covariance matrix estimation lacks rigorous finite-sample theoretical foundations, particularly under hold-out validation. Method: We derive the first closed-form analytical expression for the estimation error under the white inverse Wishart population model, integrating high-dimensional statistics, random matrix theory, and asymptotic spectral analysis. Contribution/Results: We establish that the optimal train-test ratio scales as $Theta(sqrt{p})$, where $p$ is the dimension; this scaling holds exactly for finite samples and unifies hold-out and $k$-fold cross-validation in the high-dimensional asymptotic limit—both converge to the optimal error bound of the nonlinear shrinkage estimator. Our work provides the first analytically tractable and empirically verifiable theoretical framework for cross-validation data partitioning in covariance estimation, resolving the longstanding gap in finite-sample performance characterization of cross-validation for this fundamental problem.
This study addresses the issues of information leakage and optimistic bias in performance evaluation arising from improper data splitting during machine learning model validation. Focusing on biomedical contexts, it systematically reviews validation strategies and contrasts flawed versus leakage-free designs through eight controlled experiments. Employing techniques such as nested grouped cross-validation for reproducible simulations, this work proposes deployment-oriented, scenario-specific validation guidelines. Its primary contributions include practical tools—namely decision trees, checklists, and code templates—that operationalize the core principle of aligning independent units with deployment objectives within an auditable evaluation framework. Ultimately, these contributions significantly enhance the reliability and standardization of machine learning model validation practices.
Standard leave-one-out cross-validation (LOO-CV) fails for long-range prediction evaluation in latent Gaussian models—especially those with structured random effects—because it implicitly assumes independent observations, violating the spatial, spatiotemporal, or network dependencies inherent in real-world extrapolation settings. To address this, we propose Group-Out Cross-Validation (Group-Out CV): a dependency-aware CV scheme that automatically identifies and excludes groups of statistically dependent observations, thereby emulating realistic out-of-sample prediction scenarios. We further introduce a Bayesian joint posterior correction method that adjusts predictive distributions without refitting the model. This constitutes the first CV framework explicitly designed for dependent data, featuring automatic grouping and no re-estimation. Implemented in the open-source R-INLA package, our approach demonstrably enhances the robustness and reliability of predictive assessment compared to LOO-CV, particularly in models with complex dependence structures.
This paper identifies a distributional bias induced by leave-one-out cross-validation (LOO-CV) in small-sample settings: the mean of the training set—excluding each held-out sample—is systematically negatively correlated with that sample’s label, leading to distorted model evaluation, particularly under strong regularization, where performance is systematically underestimated. To address this, the paper formally defines and quantifies the bias for the first time, and proposes ReBalanced CV—a scalable, reweighting-based cross-validation framework that calibrates training-set distribution via importance-weighted resampling. Theoretical analysis and extensive experiments on synthetic and real-world datasets—spanning logistic regression, random forests, and neural networks, and evaluating AUC-ROC and AUC-PR—demonstrate that ReBalanced CV significantly improves the accuracy of LOO-CV performance estimates, mitigates regularization bias in hyperparameter optimization, and enhances selection robustness.
Standard errors for LOO-CV predictive performance estimates in Bayesian model comparison are systematically underestimated—especially under small samples, model misspecification, or near-identical model predictions—due to persistent skewness in the LOO-CV error distribution, even asymptotically. Method: We establish the first theoretical proof that LOO-CV error skewness remains pathological in the infinite-data limit, invalidating Gaussian approximations. Building on this, we develop a rigorous yet practical uncertainty quantification framework integrating asymptotic analysis, Monte Carlo simulation, and explicit error distribution modeling. Contribution/Results: We propose a diagnosable skewness warning boundary and a robust calibration procedure that substantially improves uncertainty reliability. Our work introduces the first distribution-shape–aware diagnostic criterion for Bayesian model comparison, enabling a paradigm shift from point-estimate–based comparison to distribution-aware comparison.
研究通过在Regent Chess环境中测试,探讨了托管语言模型的行为评估差异问题,使用了复制、测量敏感性和持久性三种验证方法。
研究解决了语言模型评估不稳定的问题,通过预注册审计发现请求重复性和一致性未达预期标准,提出设计规则和报告清单以改善测量可靠性。
When data are grouped, hierarchical or multilevel models are commonly used to account for group-level variation with group-specific parameters. Leave-one-group-out cross-validation (LOGO-CV) is a suitable tool for evaluating predictive performance for new groups, providing an estimator of the expected log predictive density (elpd). Brute-force LOGO-CV requires one model refit per held-out group, often using computationally expensive inference algorithms such as MCMC. This is costly, particularly for large numbers of groups or complex model structures. Commonly used importance sampling approximations, intended to reduce this cost, tend to fail because the group-specific parameters of the held-out group must be integrated out. We identify two key challenges in LOGO-CV elpd estimation: approximating the LOGO posterior and computing the grouped marginal likelihood. We compare 11 strategies, including 5 newly proposed, to address them. Among others, we combine Pareto-smoothed importance sampling or adaptive importance sampling with integration techniques such as Laplace approximation, adaptive Gauss-Hermite quadrature, and bridge sampling. We evaluate these strategies in both simulation experiments and real-world case studies, which show that marginalising over the group-specific parameters substantially improves the reliability of the importance sampling approaches.
本文提出了一种高效计算嵌套交叉验证预测区间的方法,针对某些惩罚回归模型,只需单次模型拟合,显著减少了计算成本。
本文提出REPVIS2设计空间,通过八个维度描述复制研究与参考研究的关系,以解决复制研究设计难以描述和比较的问题。