Score
Designs and executes experimental protocols, measurement procedures, and statistical analyses to quantify and characterize a model's generalization behavior — e.g., constructing metrics and validation/training comparisons to detect overfitting, testing for memorization versus representation learning, and measuring generalization gaps. Produces evaluations of scaling and distributional effects (scaling analysis, scale generalization evaluation), interprets empirical results, and recommends interventions such as regularization or augmentation based on those analyses.
Existing evaluation methodologies for large language models (LLMs) suffer from insufficient generalization assessment, as static benchmarks fail to capture the continuously expanding capability boundaries of evolving LLMs. Method: We formally define “evaluation generalizability” and propose a four-dimensional analytical framework encompassing evaluation methodologies, datasets, evaluators, and metrics. Our approach innovatively integrates LLM-as-a-judge, dynamically updated datasets, capability-decoupled benchmark design, and a multidimensional meta-evaluation framework. Contribution: We establish a novel, capability-oriented, automated, and sustainably evolvable evaluation paradigm covering critical dimensions—including knowledge, reasoning, instruction following, multimodal understanding, and safety. Concurrently, we release an open-source, extensible GitHub “living review” repository—a community-maintained, versioned resource—to advance evaluation practice from static benchmarking toward dynamic, collaborative co-evolution.
Bayesian optimal experimental design (BOED) suffers significant generalization degradation under model misspecification and covariate shift—critical challenges in drug discovery and clinical trials. Method: We first identify and formalize a dual mechanism of “error amplification versus suppression,” enabling a decomposable theoretical framework for generalization error. Leveraging this insight, we propose a novel acquisition function that jointly ensures representativeness and error-dampening properties. Our approach integrates BOED, generalization error analysis, and covariate shift modeling, explicitly mitigating error accumulation from distributional shifts while preserving computational tractability. Contribution/Results: Experiments across diverse misspecification settings demonstrate that our method consistently outperforms standard BOED, reducing average generalization error by 27–41%. This establishes a robust experimental design paradigm for high-stakes, low-tolerance scientific decision-making.
The generalizability of machine learning experimental results—i.e., consistency across varying conditions—has long lacked rigorous, quantitative assessment due to the absence of mathematical formalization of experimental procedures. Method: This paper introduces the first principled mathematical modeling framework for ML experiments, treating them as stochastic processes and defining computable, reproducible generalizability metrics. The approach integrates probabilistic modeling, statistical inference, and experimental design theory to enable both diagnostic analysis of generalizability and estimation of minimal required sample sizes. Contribution/Results: Applied to ImageNet and GLUE benchmarks, the framework successfully identifies generalizability boundaries for several widely cited conclusions. A fully open-source Python toolkit implements end-to-end reproducibility and supports community-driven extensions. This work establishes the first verifiable, quantitative foundation for scientific rigor in ML experimentation.
Scientific machine learning experiments often suffer from distorted performance evaluations due to poor experimental design and inconsistent documentation. To address this, we propose a principled framework for ML experimentation tailored to scientific research, encompassing data preprocessing, model selection, cross-validation, and reporting—emphasizing reproducibility, fair comparison, and transparency. Our key contributions include two novel quantitative metrics: the Logarithmic Overfitting Ratio (LOR) and Composite Overfitting Score (COS), which jointly characterize overfitting severity and instability across cross-validation folds. Complementing these, we introduce standardized preprocessing protocols, rigorously defined strong baselines, and modular visualization templates for diagnostic analysis. Empirical evaluation demonstrates that our framework substantially enhances experimental rigor, reproducibility, and result credibility in scientific ML. It further enables robust performance assessment and cross-study comparability, providing systematic support for establishing reliable benchmarks.
This paper addresses the limited out-of-distribution (OOD) generalization of medical AI models in real-world clinical settings. We propose the first three-tier generalization capability scale specifically designed for medical artificial intelligence, systematically characterizing model performance under varying target-domain data and label availability—such as cross-institutional, cross-device, and cross-population scenarios—and unifying the modeling of generalization behavior across diverse deployment constraints. Grounded in theoretical analysis and empirical validation across clinical use cases, our framework enables graded assessment and informs adaptive strategy selection. It provides researchers with actionable evaluation criteria and a principled development roadmap. By bridging the gap between laboratory validation and large-scale clinical deployment, this work significantly enhances the robustness and practical applicability of medical AI models in complex, heterogeneous real-world environments.
High-complexity machine learning models lack reliable, theoretically grounded mechanisms for detecting overfitting. Method: We propose a statistical hypothesis test that operates solely on training data, dispensing with the need for an independent validation set or PAC-style uniform convergence assumptions. Our approach formalizes overfitting via empirical mean consistency and constructs a rigorous testing framework based on Hoeffding-type concentration inequalities. Contribution/Results: This is the first method to use empirical mean consistency as an overfitting criterion, enabling significance-based inference and implicit diagnosis sensitive to distributional shifts. We prove its validity under mild regularity conditions. Empirical evaluation demonstrates robust identification of overfitting transition points and latent distribution drift, substantially improving both the reliability and interpretability of model selection.
This work investigates how model width and training sample size jointly influence the generalization performance of finite-width, two-layer quadratic neural networks with ℓ² regularization under structured, finite-sample data regimes. Leveraging spectral analysis and finite-sample generalization theory, the study establishes—for the first time in finite-width networks capable of feature learning—an explicit, data-dependent expression for generalization error dominated by the spectral structure of the target function. This expression reveals power-law relationships governing how generalization error scales with network width, sample size, and regularization strength, identifies multiple scaling regimes and their phase-transition boundaries (such as the interpolation threshold), and demonstrates that the spectral structure of the data fundamentally determines the exponent of the generalization power law.
Traditional hybrid experimental designs struggle to robustly control the frequentist operating characteristics of Bayesian decisions under model misspecification and lack efficient sample size determination methods applicable to generalized posteriors. This work proposes a computationally efficient experimental design framework that requires simulations at only two sample sizes and leverages extrapolation modeling of posterior summary functions to infer performance across the entire sample size space. This approach enables identification of the minimal sample size and decision rule satisfying desired operating characteristics. It represents the first general and scalable method for sample size planning under generalized posteriors, substantially reducing computational burden while enhancing robustness to model misspecification. The method’s validity and broad applicability within Bayesian M-estimation–type experiments are demonstrated through the redesign of an adaptive clinical trial with time-to-event outcomes.
This study systematically investigates the mechanisms by which data scale, model complexity, and input modality influence the generalization performance of vision models. Within a unified experimental framework, the authors conduct controlled and large-scale ablation studies on synthetic functions and the CIFAR dataset, employing polynomial fitting, diverse CNN and Transformer architectures, and multimodal inputs—including RGB, grayscale, gradients, edges, and wavelet representations—to quantitatively compare the effects of these three core factors for the first time. The findings reveal that increasing training data consistently enhances generalization; greater model complexity yields non-monotonic improvements; removing color information substantially degrades performance; and the efficacy of explicit handcrafted priors is highly dependent on model architecture.
This study addresses the underexplored engineering challenges in existing machine learning evaluation frameworks, where operational issues and their root causes have lacked systematic investigation. To bridge this gap, the work formally establishes evaluation engineering as a distinct research direction within software engineering. Through an empirical analysis of 57 frameworks and a comprehensive categorization of 16,560 reported issues across a newly proposed five-stage workflow model, the study reveals that 41.4% of problems originate in the specification phase, while 61.7% of classified issues stem from missing functionality, inadequate documentation, and insufficient input validation. The findings yield a structured taxonomy of evaluation-related problems and provide empirical evidence to inform the design and improvement of robust evaluation systems.