Score
Designs and implements model training pipelines and optimization procedures, including data splits, loss functions, optimizers, and hyperparameter selection, to fit model parameters reliably. Builds and runs experimental evaluation protocols and metrics to measure and analyze models' generalization and behavior, compare alternatives across settings, and support model selection and validation.
This paper identifies a systemic issue in machine learning: preprocessing hyperparameters—such as missing-value imputation strategies—are frequently overlooked yet substantially bias model evaluation. Current practice often involves informal, post-hoc tuning of preprocessing steps, leading to optimistic performance estimates and irreproducible results. To address this, the authors formally distinguish and empirically analyze the coupling effects between algorithmic and preprocessing hyperparameters. Using a modular supervised learning workflow model, controlled variable experiments, replication of canonical case studies, and bias diagnostics, they quantify the resulting optimistic bias. Key contributions include: (1) establishing preprocessing hyperparameters as equally critical as algorithmic ones; (2) proposing formal modeling principles to eliminate informal preprocessing tuning; and (3) delivering actionable reporting guidelines for ML practitioners, thereby significantly enhancing model credibility and reproducibility.
This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.
To address the challenge of efficiently optimizing high-dimensional configuration spaces in large-scale machine learning training, this paper proposes a scalable meta-gradient computation algorithm and the Smooth Model Training (SMT) framework—enabling, for the first time, end-to-end, differentiable joint optimization of training strategies. Methodologically, it integrates reverse-mode automatic differentiation through training loops, smooth modeling of training trajectories, and meta-gradient descent (MGD) to jointly optimize data selection, poisoning-resilient strategies, and learning rate scheduling. Key contributions are: (1) a breakthrough in scalable meta-gradient computation for large-scale training; and (2) the SMT framework, which ensures stability and convergence of MGD under realistic dynamic training conditions. Experiments demonstrate that the proposed data selection method significantly outperforms existing approaches; robustness against accuracy-degrading data poisoning attacks improves by an order of magnitude; and the fully automated learning rate scheduler matches or exceeds hand-crafted designs in performance.
This work addresses the mismatch between conventional machine learning practices and the specific performance requirements of clinical tasks in healthcare settings. Traditional approaches rely on differentiable validation losses for model optimization, which often fail to align with clinically meaningful outcomes. To bridge this gap, the paper proposes replacing standard loss functions with non-differentiable yet clinically interpretable custom metrics to guide critical optimization decisions—such as hyperparameter selection and training termination—thereby redefining the model validation pipeline. In two controlled experiments, models optimized using this framework demonstrated significantly superior performance on key clinical tasks compared to those guided by conventional differentiable validation losses. This approach overcomes the inherent limitation of relying solely on differentiable objectives and better aligns medical AI development with real-world clinical goals.
This paper addresses the inefficiency and lack of scalability of manual hyperparameter tuning in large-scale machine learning. It systematically surveys hyperparameter optimization (HPO), unifying and classifying five mainstream paradigms: random/low-discrepancy search, bandit-based methods, Bayesian optimization, population-based (evolutionary) algorithms, and gradient-based differentiable optimization. The survey further extends to emerging settings—including online HPO, constrained HPO, and multi-objective HPO. Crucially, the work establishes novel theoretical connections between HPO and meta-learning as well as neural architecture search, yielding a comprehensive knowledge framework that articulates methodological principles, applicability boundaries, and inherent limitations. By clarifying the technical evolution and identifying key open challenges, this study provides a theoretically grounded yet practically actionable foundation for automated machine learning.
This paper addresses the challenge of evaluating the generalizability of data-driven models to target populations. Methodologically, it integrates statistical learning theory with empirical validation paradigms to propose the first systematic, domain-agnostic framework for model validation. The framework establishes three core principles: (1) sufficiency of validation strategies, (2) mandatory disclosure of limitations, and (3) design of performance metrics enabling cross-model comparability. Its key contribution lies in unifying transparency, reproducibility, and methodological rigor within a single, generalizable validation standard—thereby overcoming longstanding fragmentation and inconsistent reporting in current validation practices. Applicable to diverse data-driven models—including machine learning and statistical prediction models—the framework substantially enhances validation reliability, result comparability across studies, and cross-study reproducibility. It provides a foundational methodological basis for clinical deployment and regulatory evaluation of predictive models.
This work addresses the challenge of high online tuning costs and the absence of simulators in real-world reinforcement learning systems by proposing a method that constructs a calibrated model from offline data to approximate environment dynamics, enabling offline hyperparameter selection. The approach is applied for the first time to a municipal water treatment plant, employing a k-nearest neighbors model with Laplacian distance for nexting prediction. Evaluated on high-dimensional, non-stationary data over annual timescales, the method demonstrates strong scalability and robustness to distributional shifts. Experimental results show that the calibrated model generates realistic long-horizon trajectories, accurately reproduces hyperparameter sensitivity trends, and effectively supports fine-grained tuning of the agent’s learning rate.
Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.
This paper addresses the challenge of multi-model-class evaluation by proposing the Model Class Selection (MCS) framework, which identifies the collection of model classes each containing at least one optimal model—thereby enabling formal comparison of performance equivalence across model classes of differing complexity (e.g., interpretable vs. black-box models). MCS generalizes conventional model selection and Model Set Selection (MSS) by integrating likelihood maximization and risk minimization criteria via a data-splitting strategy under mild assumptions. Theoretical analysis establishes its statistical validity. Empirical evaluation—including simulations and real-data experiments—demonstrates that MCS robustly identifies simple, interpretable model classes whose predictive performance matches that of complex models. By bridging interpretability and performance assessment, MCS introduces a novel paradigm and practical tool for explainable AI research.