dimension-wise generalization testing

Design and implement leave-one-dimension-out evaluation protocols that systematically hold out a single configuration or input dimension to test model and cross-model generalization. Build analyses and metrics that measure predictive accuracy across regimes, identify architectural boundaries where performance degrades, and map regimes to surrogate-model reliability.

dimension-wisegeneralizationtesting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.09
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Fixed-size benchmarking in model evaluation often fails to balance efficiency, statistical reliability, and diverse objectives, leading to either excessive resource consumption or unreliable results. This work proposes the first adaptive framework that integrates sequential testing into AI model evaluation, dynamically allocating evaluation data based on stopping criteria tailored for model ranking and selection tasks. By combining sequential hypothesis testing, minimum detectable effect analysis, and diminishing returns detection, the method achieves substantial gains in efficiency without compromising rigor. Empirical validation on the Open VLM Leaderboard demonstrates an 80% reduction in computational cost while maintaining a confidence interval width of 2.5 points, significantly enhancing both the practicality and scalability of model evaluation.

computational costfixed-size benchmarksmodel evaluation

A New Flexible Train-Test Split Algorithm, an approach for choosing among the Hold-out, K-fold cross-validation, and Hold-out iteration

Jan 11, 2025
ZB
Zahra Bami
🏛️ University of Turin | University of Social Welfare and Rehabilitation Science | Macquarie University

Model evaluation strategies significantly impact reported accuracy, yet systematic comparisons of cross-validation methods across diverse datasets and algorithms remain limited. Method: This study conducts a comprehensive empirical analysis of hold-out, iterative hold-out, and K-fold cross-validation across two real-world datasets (Framingham Heart Study and COVID-19) and three classifiers (decision tree, naive Bayes, k-nearest neighbors). A parameter sensitivity analysis investigates interactions among test set proportion, random seed, and fold count (K). Contribution/Results: We find that a 10% hold-out split consistently yields higher accuracy than the conventional 20% split across most configurations; iterative hold-out substantially reduces variance; and no universally optimal K exists—its efficacy depends critically on dataset size, feature distribution, and learning algorithm. The study demonstrates that validation strategy selection must jointly account for data characteristics and model properties, providing a reproducible evaluation framework and empirically grounded guidelines for robust model assessment.

Cross-validation methodsMachine Learning Model EvaluationParameter Configuration

Distributional bias compromises leave-one-out cross-validation

Jun 03, 2024
GI
George I. Austin
🏛️ Columbia University | Columbia University Irving Medical Center

This paper identifies a distributional bias induced by leave-one-out cross-validation (LOO-CV) in small-sample settings: the mean of the training set—excluding each held-out sample—is systematically negatively correlated with that sample’s label, leading to distorted model evaluation, particularly under strong regularization, where performance is systematically underestimated. To address this, the paper formally defines and quantifies the bias for the first time, and proposes ReBalanced CV—a scalable, reweighting-based cross-validation framework that calibrates training-set distribution via importance-weighted resampling. Theoretical analysis and extensive experiments on synthetic and real-world datasets—spanning logistic regression, random forests, and neural networks, and evaluating AUC-ROC and AUC-PR—demonstrate that ReBalanced CV significantly improves the accuracy of LOO-CV performance estimates, mitigates regularization bias in hyperparameter optimization, and enhances selection robustness.

Distributional bias affects leave-one-out cross-validation accuracyNegative correlation between training and test labels skews evaluationProposed rebalanced cross-validation corrects bias in classification and regression

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities

Dec 09, 2024
AG
Adhiraj Ghosh
🏛️ University of Tübingen | Open-Ψ (Open-Sci) Collective | University of Cambridge

Traditional static test sets inadequately evaluate foundation models’ diverse capabilities in open-ended scenarios. To address this, we propose ONEBench—a dynamic, extensible benchmarking paradigm that enables on-demand generation of customized evaluation suites targeting open capabilities, framing model assessment as a collective selection and aggregation process over sample-level tests. Our key contributions include: (1) the first unified, open-ended, and evolvable evaluation framework operating at the sample level; (2) a sparse measurement aggregation algorithm, a progressive sample pool construction mechanism, and a cross-modal unified interface (ONEBench-LLM/LMM); and (3) a robustness-aware scoring model with theoretical guarantees on identifiability and fast convergence. Experiments show that ONEBench achieves ranking stability >0.98 under 95% measurement sparsity, reduces evaluation cost by 20×, and attains >0.98 correlation with mean-score rankings on homogeneous data—enabling unified, efficient, and reliable assessment of both language and multimodal models.

Aggregating diverse metrics into reliable model scoresEvaluating open-ended capabilities of foundation modelsReducing evaluation cost while maintaining accuracy

This paper addresses the challenge of quantifying individual model contributions in multi-model ensemble forecasting. We propose an interpretable attribution framework grounded in Shapley values from cooperative game theory—the first application of Shapley values to ensemble importance assessment. To ensure scalability and theoretical rigor, we introduce two efficient algorithms: Leave-One-Model-Out (LOMO) and Leave-All-Subsets-of-Models-Out (LASMO). By integrating error similarity analysis and Monte Carlo approximation, we significantly reduce computational complexity. Evaluated on the US COVID-19 mortality prediction task, our method identifies models with low standalone accuracy but high collaborative value—revealing complementary and redundant interactions among models that conventional accuracy metrics fail to capture. The framework advances ensemble interpretability and informs principled model selection, establishing a new paradigm for explainable ensemble learning.

Measuring individual model importance in ensemble forecasting accuracyProposing practical methods to assess model contribution to ensemble performanceRevealing unique model features beyond standard accuracy metrics

Latest Papers

What's happening recently
View more

This work addresses the high cost of ground-truth evaluation in chemical and materials design, where existing machine learning surrogate models often lack reliability guarantees. Departing from conventional reliance on prediction accuracy metrics such as R²—which can paradoxically increase the risk of worst-case selections—the study proposes “rank preservation” as a core criterion for surrogate validation. It formally introduces the concept of “selection tax” and derives its theoretical upper and lower bounds. A safety certification framework for surrogates is established through selection-aware auditing, rank correlation analysis, and multi-task ground-truth validation. Experiments demonstrate that the proposed audit statistics achieve Spearman correlations of 0.80–0.99 with actual search performance, substantially outperforming R² (as low as 0.33). Certified screening strategies based on this framework reduce evaluation costs by up to 25-fold.

experimental replacementmodel validationselection bias

This study addresses the critical issue of model instability in software engineering optimization, which leads to substantial variability across repeated experiments and undermines both credibility and practical utility. Rather than treating instability as mere random noise, this work conceptualizes it as a quantifiable and manageable property that should be integrated into standard evaluation frameworks. By systematically modulating label usage, model complexity, and partition scoring strategies—combined with multi-objective optimization, causal intervention, data locality analysis, and model calibration—the proposed approach significantly enhances result consistency. Empirical evaluation demonstrates that the optimized configuration reduces the standard deviation of error by 22% on average and outperforms default settings in 119 out of 127 datasets, achieving a 4.8-fold improvement in result consistency.

model instabilitymulti-objective optimizationreproducibility

This work addresses the challenge of certifying performance attributes that emerge as user concerns after deployment but were not considered during the design phase in data-driven control. To this end, the paper proposes a two-layer adaptability framework that extends the scenario approach by introducing a post-design adaptability concept, enabling reliable certification without requiring additional test data. It is the first to formally incorporate user-specified a posteriori performance properties into the scenario optimization framework, deriving computable, distribution-free upper and lower bounds on the violation risk. Moreover, the method allows full reconstruction of the performance metric’s distribution from existing data. Experimental validation on H₂ control and pole placement problems demonstrates that the approach effectively certifies a posteriori properties and accurately infers critical performance distributions, offering both theoretical rigor and practical utility.

distribution-free boundsgeneralizationpost-design certification

该研究通过消除几何学框架探讨局部最优对象能否由共享部署规则实现,分析信息、架构等因素对缺陷修复的影响。

defect visibilityElimination Geometryinformation loss

Hot Scholars

XZ

Xiangliang Zhang

Leonard C. Bettex Collegiate Professor, Computer Science and Engineering, University of Notre Dame
Machine LearningAI for Science
TL

Tal Linzen

New York University
Language modelsComputational linguisticsNatural language processingCognitive science
YZ

Yingqian Zhang

Associate Professor of AI for Decision-Making, Eindhoven University of Technology
Artificial IntelligenceData-Driven OptimizationDeep RLSocial-aware Algorithms
YW

Yaoxin Wu

Eindhoven University of Technology
Deep learningCombinatorial optimizationInteger programmingMulti-objective optimization
YS

Yangqiu Song

HKUST
Artificial IntelligenceData MiningNatural Language ProcessingKnowledge Graphs