model selection

Designs, builds, and analyses procedures, criteria, and workflows for choosing among candidate models or model families and their hyperparameters to meet specified performance and deployment requirements. This includes creating evaluation pipelines and validation schemes (e.g., cross‑validation), information criteria and hypothesis tests, selection algorithms and metrics, and analyses of trade‑offs such as complexity vs. generalization, robustness, and computational or resource constraints.

modelselection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$193K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Ensuring Reliability via Hyperparameter Selection: Review and Advances

Feb 06, 2025
AF
Amirmohammad Farzaneh
🏛️ King’s College London

To address reliability assurance of pretrained foundation models in engineering deployments—particularly in communication systems—this paper formulates hyperparameter selection as a multiple hypothesis testing problem and, for the first time, derives verifiable statistical upper bounds on population risk within the Learn-Then-Test (LTT) framework. Methodologically, we introduce four key innovations: (i) multi-objective hyperparameter optimization, (ii) Bayesian integration of domain-specific prior knowledge, (iii) graph-based modeling of task-dependent structural dependencies, and (iv) an adaptive selection mechanism tailored to heterogeneous risk metrics. The proposed framework uniformly supports diverse risk measures—including bit error rate and calibration error—and accommodates varying strengths of statistical guarantees, such as false discovery rate (FDR) control and high-confidence upper bounds. Empirical evaluation on communication systems demonstrates substantial improvements in deployment reliability, with analytically verifiable risk bounds.

Extensions for engineering-relevant scenariosHyperparameter selection in AI modelsStatistical guarantees on risk measures

This work addresses the lack of statistical reliability guarantees in existing hyperparameter selection methods—such as grid search and Bayesian optimization—with respect to critical metrics like risk and safety. Building upon the learn-then-test (LTT) paradigm, the paper introduces a unified statistical framework that formulates hyperparameter selection as a multiple hypothesis testing problem, accommodating user-specified constraints on average risk, quantile risk, or information-theoretic measures. Leveraging tools from statistical inference—including p-values, e-values, and concentration inequalities—the method derives explicit, finite-sample bounds on error probabilities from first principles. This approach enables theoretically grounded validation and selection of hyperparameters, substantially enhancing the reliability and safety of AI systems in real-world deployment scenarios.

hyperparameter selectionmultiple hypothesis testingreliability

Beyond algorithm hyperparameters: on preprocessing hyperparameters and associated pitfalls in machine learning applications

Dec 04, 2024
CS
Christina Sauer
🏛️ LMU Munich | Munich Center for Machine Learning | Medical University of Vienna

This paper identifies a systemic issue in machine learning: preprocessing hyperparameters—such as missing-value imputation strategies—are frequently overlooked yet substantially bias model evaluation. Current practice often involves informal, post-hoc tuning of preprocessing steps, leading to optimistic performance estimates and irreproducible results. To address this, the authors formally distinguish and empirically analyze the coupling effects between algorithmic and preprocessing hyperparameters. Using a modular supervised learning workflow model, controlled variable experiments, replication of canonical case studies, and bias diagnostics, they quantify the resulting optimistic bias. Key contributions include: (1) establishing preprocessing hyperparameters as equally critical as algorithmic ones; (2) proposing formal modeling principles to eliminate informal preprocessing tuning; and (3) delivering actionable reporting guidelines for ML practitioners, thereby significantly enhancing model credibility and reproducibility.

Addresses overlooked preprocessing hyperparameters in ML model tuningAims to improve predictive modeling quality and reportingHighlights pitfalls in informal preprocessing optimization practices

This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.

compression and accelerationconstraint-drivendeployment constraints

This paper addresses the inefficiency and lack of scalability of manual hyperparameter tuning in large-scale machine learning. It systematically surveys hyperparameter optimization (HPO), unifying and classifying five mainstream paradigms: random/low-discrepancy search, bandit-based methods, Bayesian optimization, population-based (evolutionary) algorithms, and gradient-based differentiable optimization. The survey further extends to emerging settings—including online HPO, constrained HPO, and multi-objective HPO. Crucially, the work establishes novel theoretical connections between HPO and meta-learning as well as neural architecture search, yielding a comprehensive knowledge framework that articulates methodological principles, applicability boundaries, and inherent limitations. By clarifying the technical evolution and identifying key open challenges, this study provides a theoretically grounded yet practically actionable foundation for automated machine learning.

Addressing challenges in online, constrained, and multi-objective hyperparameter tuningAutomating hyperparameter search to improve machine learning efficiencyComparing state-of-the-art hyperparameter optimization techniques and methods

Latest Papers

What's happening recently
View more

Evidence-Driven Decision Support for AI Model Selection in Research Software Engineering

Dec 12, 2025
AJ
Alireza Joonbakhsh
🏛️ Shiraz University | Wageningen University & Research

Rapid AI model evolution has led to ad hoc, non-reproducible model selection in scientific software engineering, severely undermining reproducibility and transparency. To address this, we propose ModelSelect—the first evidence-driven framework that formalizes AI model selection as a multi-criteria decision-making (MCDM) problem, integrating automated metadata harvesting, a structured knowledge graph, and research-context-aware decision modeling. Its key contributions are: (1) establishing the first MCDM modeling paradigm tailored to scientific practice; (2) end-to-end integration of technical metrics and domain semantics, significantly enhancing recommendation interpretability and consistency; and (3) empirical validation across 50 real-world research scenarios, achieving 96.2% coverage, substantially higher rationale alignment than baselines, and superior traceability, cross-scenario robustness, and transparency.

Addresses ad hoc model selection undermining reproducibility and transparencyFrames model selection as a Multi-Criteria Decision-Making problemStructured evidence-driven AI model selection for research software engineering

This work addresses the inefficiencies in large-scale recommendation systems caused by maintaining separate models for different scenarios and objectives, which hinders development velocity and delays technology adoption. To overcome this, the authors propose the Standardized Model Template (SMT) framework, which leverages composable, standardized machine learning components to enable “design once, deploy everywhere,” uniformly accommodating diverse data distributions and optimization objectives. By decoupling model architecture from scenario-specific configurations, SMT reduces the complexity of technology deployment from O(n·2ᵏ) to O(n+k), breaking away from the conventional “one objective, one model” paradigm. Empirical evaluation on Meta’s ad ranking system demonstrates that SMT improves average cross-entropy by 0.63%, reduces engineering time per model iteration by 92%, and increases the throughput of technology-model pair adoption by 6.3×.

computational advertisinglarge-scale ML ecosystemsML technique propagation

Surrogate Benchmarks for Model Merging Optimization

Sep 02, 2025
RA
Rio Akizuki
🏛️ Yokohama National University

Hyperparameter optimization for large language model (LLM) ensembling incurs prohibitive computational costs and hinders efficient evaluation of optimization algorithms. Method: We propose a lightweight surrogate benchmark that defines a multidimensional hyperparameter search space, collects a small number of real ensembling experiments, and constructs a regression-based surrogate model to accurately predict ensemble performance across hyperparameter configurations. Contribution/Results: Compared to direct tuning, our surrogate reduces evaluation cost by over two orders of magnitude while faithfully reproducing the convergence behavior and relative ranking of optimization algorithms. Experiments across multiple LLM ensembling tasks demonstrate high predictive accuracy (average MAE < 0.8%) and strong generalization. The benchmark provides an efficient, reproducible, and low-cost standardized testbed for developing, comparing, and deploying hyperparameter optimization algorithms for LLM ensembling.

Developing surrogate benchmarks for performance predictionOptimizing hyperparameters for model merging techniquesReducing computational costs in merging large language models

This work addresses the lack of systematic, auditable decision support in foundational model selection, a process often driven by popularity rather than functional capabilities, operational constraints, or community-based quality assessments. To remedy this, we propose HugSelect—the first interpretable multi-criteria decision framework that formalizes foundational model selection as a traceable software component selection task. HugSelect integrates model metadata, functional attributes, and community-perceived quality into a unified knowledge base and a decomposable weighted scoring system. Evaluated on 71,274 models, our approach achieves an F1 score of 0.801 in functional feature extraction and 0.84 accuracy in quality mapping. It attains model-level Coverage@10 of 0.61 and family-level Coverage@10 of 0.91, matching the recommendation performance of leading commercial systems while offering transparent, interpretable results that elicit positive user feedback.

explainable AIfoundation-model selectionmodel hubs

This work addresses a critical gap in the testing of machine learning (ML) components, where existing approaches predominantly focus on model performance while neglecting system-level quality attributes such as throughput, resource consumption, and robustness—often leading to integration failures. To bridge this gap, the paper proposes the first standalone quality model specifically tailored for ML components. Grounded in the ISO/IEC 25010 quality standard framework and informed by requirements engineering and software quality modeling techniques, the model systematically decouples and structures key quality attributes of ML components, thereby addressing the lack of component-level applicability in ISO/IEC 25059. It provides developers and stakeholders with a unified terminology to prioritize testing efforts. The model’s effectiveness has been validated through user studies and has been integrated into an open-source ML testing tool, enabling practical deployment.

ISO 25059machine learning componentsquality model

Hot Scholars

PL

Peng Liang

School of Computer Science, Wuhan University
Software EngineeringSoftware ArchitectureEmpirical Software Engineering
RC

Rohitash Chandra

UNSW
Bayesian deep learningNeuroevolutionClimate ExtremesLanguage Models
MS

Mark Stamp

Professor of Computer Science, San Jose State University
information securitycryptographymachine learning
SB

Siming Bayer

Researcher, Pattern Recognition Lab, Friedrich-Alexander University
Medical Image ProcessingComputer Guided InterventionMachine Learning
PA

Paula Andrea Pérez-Toro

Friedrich-Alexander-Universität Erlangen-Nürnberg; Universidad de Antioquia
Machine LearningSpeech AnalysisGait AnalysisNatural Language Processing