Score
Designs and implements algorithms, pipelines, and evaluation frameworks that select or recommend predictive models based on dataset characteristics, label structure, and resource or class-count constraints. This includes defining selection criteria and metrics (statistical, automated, and task-guided), building adaptive/dynamic and sequential model-selection procedures for few-class or data-scarce settings, and analyzing dataset-to-model links (e.g., feature recoverability, feature–label relationships, or embedding rankings) to guide recommendations.
The algorithm selection and parameterization (ASP) domain lacks systematic surveys and empirical evaluations. Method: We propose the first standardized, meta-learning–driven ASP framework, built upon the largest ASP benchmark knowledge base to date—comprising 400 datasets and 4 million pre-trained models—and conduct large-scale comparative experiments across eight mainstream classifiers under diverse scenarios. Our evaluation integrates empirical performance modeling (EPM), feature engineering, and statistical significance testing to quantify accuracy, generalizability, and computational efficiency. Contribution/Results: This work delivers the first critical survey balancing methodological rigor with empirical breadth; reveals performance boundaries and applicability conditions of state-of-the-art ASP methods; and establishes a reproducible benchmark and practical selection guide for AutoML research and deployment.
Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.
This paper addresses the lack of standardized evaluation criteria for machine learning algorithm selection in healthcare, telecommunications, and marketing. We propose a multidimensional, automated model selection framework that jointly optimizes predictive performance—measured by accuracy, precision, and recall—and model complexity, as quantified by the Akaike Information Criterion (AIC). The framework is designed to be domain-agnostic, supporting seamless adaptation across eager, lazy, and hybrid learning paradigms. Evaluated on eight real-world datasets—including cardiovascular disease prediction and fetal health classification—the method consistently identifies optimal models, achieving statistically significant improvements in both predictive accuracy and generalization performance. Crucially, it delivers interpretable and reusable model recommendations tailored to mission-critical applications, thereby bridging the gap between theoretical model selection and practical deployment.
Machine learning model selection lacks formalized methodologies, making it difficult to systematically characterize contextual factors—such as data characteristics and prediction tasks—and their interactions, resulting in opaque, non-adaptive decisions. This paper introduces, for the first time, software product line (SPL) principles into ML model selection, proposing a variability-aware algorithm selection framework. It constructs a configurable feature model that explicitly captures commonalities and variabilities among contextual factors—including dataset size, feature dimensionality, and task type—as well as their logical dependencies. By integrating scikit-learn’s heuristic rules with an instantiation framework, the approach enables interpretable, adaptive, and transparent model recommendations. An empirical case study demonstrates that the method significantly outperforms existing strategies in accuracy, interpretability, and contextual adaptability.
This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.
Communication barriers between data scientists and domain experts arise from oversimplified, accuracy-centric model performance reporting, hindering shared understanding of model limitations and contextual applicability. Method: We propose a visualization-mediated model explanation framework grounded in human-computer interaction principles, participatory design, and visual narrative techniques. This yields the first domain-expert-oriented model communication guideline—emphasizing risk, trade-offs, and situational appropriateness rather than isolated metrics like accuracy. An iterative empirical study was conducted using regression models, incorporating structured expert feedback for evaluation. Contribution/Results: The framework significantly improves domain experts’ ability to identify model limitations, recognize inherent trade-offs, and proactively make context-driven adoption decisions. Its core innovation lies in repositioning visualization as an interdisciplinary consensus-building medium—shifting the paradigm from “metric reporting” to “collaborative understanding.”
This study addresses the limitations of rule-based model selection in time series forecasting, which often fails due to variations in data-generating mechanisms. The authors propose a descriptor framework grounded in measurable characteristics—such as trend strength, seasonality, noise level, and temporal dependence—and systematically evaluate the efficacy of static selection rules across diverse real-world datasets. Their analysis reveals, for the first time, that static descriptors are insufficient for reliably predicting model performance: model selection proves highly context-dependent and unstable. Notably, under noisy conditions or mixed generative mechanisms, recommended models frequently diverge substantially from the true best-performing ones, and model rankings exhibit pronounced sensitivity to both data characteristics and forecast horizons.
This work addresses the high cost of machine learning benchmarking by proposing a systematic framework to efficiently select small, representative subsets of datasets while preserving model ranking stability. The study presents the first comprehensive evaluation of various dataset selection strategies—including clustering, A/D-optimal experimental designs, random baselines, and a greedy farthest-first (FAFI) approach—on rank fidelity. It derives a theoretical upper bound on Spearman rank correlation error for FAFI and integrates bootstrap aggregation to yield statistically rigorous confidence intervals for comparing strategy performance. Empirical results demonstrate that as few as five datasets suffice to achieve 0.95 rank correlation in time series classification, significantly outperforming random selection in NLP tasks, though gains are limited in recommendation systems.
Rapid AI model evolution has led to ad hoc, non-reproducible model selection in scientific software engineering, severely undermining reproducibility and transparency. To address this, we propose ModelSelect—the first evidence-driven framework that formalizes AI model selection as a multi-criteria decision-making (MCDM) problem, integrating automated metadata harvesting, a structured knowledge graph, and research-context-aware decision modeling. Its key contributions are: (1) establishing the first MCDM modeling paradigm tailored to scientific practice; (2) end-to-end integration of technical metrics and domain semantics, significantly enhancing recommendation interpretability and consistency; and (3) empirical validation across 50 real-world research scenarios, achieving 96.2% coverage, substantially higher rationale alignment than baselines, and superior traceability, cross-scenario robustness, and transparency.
This work addresses the inefficiencies in large-scale recommendation systems caused by maintaining separate models for different scenarios and objectives, which hinders development velocity and delays technology adoption. To overcome this, the authors propose the Standardized Model Template (SMT) framework, which leverages composable, standardized machine learning components to enable “design once, deploy everywhere,” uniformly accommodating diverse data distributions and optimization objectives. By decoupling model architecture from scenario-specific configurations, SMT reduces the complexity of technology deployment from O(n·2ᵏ) to O(n+k), breaking away from the conventional “one objective, one model” paradigm. Empirical evaluation on Meta’s ad ranking system demonstrates that SMT improves average cross-entropy by 0.63%, reduces engineering time per model iteration by 92%, and increases the throughput of technology-model pair adoption by 6.3×.