Score
Design and implement leave-one-dimension-out evaluation protocols that systematically hold out a single configuration or input dimension to test model and cross-model generalization. Build analyses and metrics that measure predictive accuracy across regimes, identify architectural boundaries where performance degrades, and map regimes to surrogate-model reliability.
Fixed-size benchmarking in model evaluation often fails to balance efficiency, statistical reliability, and diverse objectives, leading to either excessive resource consumption or unreliable results. This work proposes the first adaptive framework that integrates sequential testing into AI model evaluation, dynamically allocating evaluation data based on stopping criteria tailored for model ranking and selection tasks. By combining sequential hypothesis testing, minimum detectable effect analysis, and diminishing returns detection, the method achieves substantial gains in efficiency without compromising rigor. Empirical validation on the Open VLM Leaderboard demonstrates an 80% reduction in computational cost while maintaining a confidence interval width of 2.5 points, significantly enhancing both the practicality and scalability of model evaluation.
Model evaluation strategies significantly impact reported accuracy, yet systematic comparisons of cross-validation methods across diverse datasets and algorithms remain limited. Method: This study conducts a comprehensive empirical analysis of hold-out, iterative hold-out, and K-fold cross-validation across two real-world datasets (Framingham Heart Study and COVID-19) and three classifiers (decision tree, naive Bayes, k-nearest neighbors). A parameter sensitivity analysis investigates interactions among test set proportion, random seed, and fold count (K). Contribution/Results: We find that a 10% hold-out split consistently yields higher accuracy than the conventional 20% split across most configurations; iterative hold-out substantially reduces variance; and no universally optimal K exists—its efficacy depends critically on dataset size, feature distribution, and learning algorithm. The study demonstrates that validation strategy selection must jointly account for data characteristics and model properties, providing a reproducible evaluation framework and empirically grounded guidelines for robust model assessment.
This paper identifies a distributional bias induced by leave-one-out cross-validation (LOO-CV) in small-sample settings: the mean of the training set—excluding each held-out sample—is systematically negatively correlated with that sample’s label, leading to distorted model evaluation, particularly under strong regularization, where performance is systematically underestimated. To address this, the paper formally defines and quantifies the bias for the first time, and proposes ReBalanced CV—a scalable, reweighting-based cross-validation framework that calibrates training-set distribution via importance-weighted resampling. Theoretical analysis and extensive experiments on synthetic and real-world datasets—spanning logistic regression, random forests, and neural networks, and evaluating AUC-ROC and AUC-PR—demonstrate that ReBalanced CV significantly improves the accuracy of LOO-CV performance estimates, mitigates regularization bias in hyperparameter optimization, and enhances selection robustness.
Traditional static test sets inadequately evaluate foundation models’ diverse capabilities in open-ended scenarios. To address this, we propose ONEBench—a dynamic, extensible benchmarking paradigm that enables on-demand generation of customized evaluation suites targeting open capabilities, framing model assessment as a collective selection and aggregation process over sample-level tests. Our key contributions include: (1) the first unified, open-ended, and evolvable evaluation framework operating at the sample level; (2) a sparse measurement aggregation algorithm, a progressive sample pool construction mechanism, and a cross-modal unified interface (ONEBench-LLM/LMM); and (3) a robustness-aware scoring model with theoretical guarantees on identifiability and fast convergence. Experiments show that ONEBench achieves ranking stability >0.98 under 95% measurement sparsity, reduces evaluation cost by 20×, and attains >0.98 correlation with mean-score rankings on homogeneous data—enabling unified, efficient, and reliable assessment of both language and multimodal models.
This paper addresses the challenge of quantifying individual model contributions in multi-model ensemble forecasting. We propose an interpretable attribution framework grounded in Shapley values from cooperative game theory—the first application of Shapley values to ensemble importance assessment. To ensure scalability and theoretical rigor, we introduce two efficient algorithms: Leave-One-Model-Out (LOMO) and Leave-All-Subsets-of-Models-Out (LASMO). By integrating error similarity analysis and Monte Carlo approximation, we significantly reduce computational complexity. Evaluated on the US COVID-19 mortality prediction task, our method identifies models with low standalone accuracy but high collaborative value—revealing complementary and redundant interactions among models that conventional accuracy metrics fail to capture. The framework advances ensemble interpretability and informs principled model selection, establishing a new paradigm for explainable ensemble learning.
This work addresses the high cost of ground-truth evaluation in chemical and materials design, where existing machine learning surrogate models often lack reliability guarantees. Departing from conventional reliance on prediction accuracy metrics such as R²—which can paradoxically increase the risk of worst-case selections—the study proposes “rank preservation” as a core criterion for surrogate validation. It formally introduces the concept of “selection tax” and derives its theoretical upper and lower bounds. A safety certification framework for surrogates is established through selection-aware auditing, rank correlation analysis, and multi-task ground-truth validation. Experiments demonstrate that the proposed audit statistics achieve Spearman correlations of 0.80–0.99 with actual search performance, substantially outperforming R² (as low as 0.33). Certified screening strategies based on this framework reduce evaluation costs by up to 25-fold.
This study addresses the critical issue of model instability in software engineering optimization, which leads to substantial variability across repeated experiments and undermines both credibility and practical utility. Rather than treating instability as mere random noise, this work conceptualizes it as a quantifiable and manageable property that should be integrated into standard evaluation frameworks. By systematically modulating label usage, model complexity, and partition scoring strategies—combined with multi-objective optimization, causal intervention, data locality analysis, and model calibration—the proposed approach significantly enhances result consistency. Empirical evaluation demonstrates that the optimized configuration reduces the standard deviation of error by 22% on average and outperforms default settings in 119 out of 127 datasets, achieving a 4.8-fold improvement in result consistency.
This work addresses the challenge of certifying performance attributes that emerge as user concerns after deployment but were not considered during the design phase in data-driven control. To this end, the paper proposes a two-layer adaptability framework that extends the scenario approach by introducing a post-design adaptability concept, enabling reliable certification without requiring additional test data. It is the first to formally incorporate user-specified a posteriori performance properties into the scenario optimization framework, deriving computable, distribution-free upper and lower bounds on the violation risk. Moreover, the method allows full reconstruction of the performance metric’s distribution from existing data. Experimental validation on H₂ control and pole placement problems demonstrates that the approach effectively certifies a posteriori properties and accurately infers critical performance distributions, offering both theoretical rigor and practical utility.
该研究通过消除几何学框架探讨局部最优对象能否由共享部署规则实现,分析信息、架构等因素对缺陷修复的影响。
论文针对代理基准测试中的双重测量混淆问题,通过将关键决策转移给模型、使用基于真实值的评分及报告更全面的可靠性指标来解决。