Score
Designs and implements systems that automatically generate candidate model variants, run evaluations against baselines, and select best-performing architectures or configurations. This work includes building search strategies, orchestration of training and evaluation pipelines, performance metrics and comparison logic to automate model interventions and final model selection.
The algorithm selection and parameterization (ASP) domain lacks systematic surveys and empirical evaluations. Method: We propose the first standardized, meta-learning–driven ASP framework, built upon the largest ASP benchmark knowledge base to date—comprising 400 datasets and 4 million pre-trained models—and conduct large-scale comparative experiments across eight mainstream classifiers under diverse scenarios. Our evaluation integrates empirical performance modeling (EPM), feature engineering, and statistical significance testing to quantify accuracy, generalizability, and computational efficiency. Contribution/Results: This work delivers the first critical survey balancing methodological rigor with empirical breadth; reveals performance boundaries and applicability conditions of state-of-the-art ASP methods; and establishes a reproducible benchmark and practical selection guide for AutoML research and deployment.
This work addresses the lack of structured, verifiable, and reusable decision mechanisms in existing automated machine learning approaches for model selection. It proposes a semantic task profiling–based structured agent framework that leverages retrieval-augmented generation of historical cases and code modules to construct an intermediate representation blueprint encompassing modeling components, composition logic, and execution constraints. By integrating code execution feedback with a failure-aware reinforcement learning strategy, the framework enables memory-driven, traceable, multi-stage search optimization. Evaluated on financial time-series forecasting and generation tasks, the method significantly outperforms both conventional AutoML systems and current agent-based baselines, achieving consistent improvements in task performance, execution success rate, and decision interpretability.
This work addresses the inefficiencies in large-scale recommendation systems caused by maintaining separate models for different scenarios and objectives, which hinders development velocity and delays technology adoption. To overcome this, the authors propose the Standardized Model Template (SMT) framework, which leverages composable, standardized machine learning components to enable “design once, deploy everywhere,” uniformly accommodating diverse data distributions and optimization objectives. By decoupling model architecture from scenario-specific configurations, SMT reduces the complexity of technology deployment from O(n·2ᵏ) to O(n+k), breaking away from the conventional “one objective, one model” paradigm. Empirical evaluation on Meta’s ad ranking system demonstrates that SMT improves average cross-entropy by 0.63%, reduces engineering time per model iteration by 92%, and increases the throughput of technology-model pair adoption by 6.3×.
Automated, systematic discovery of latent capabilities and risks in foundation models (e.g., GPT, Claude, Llama) remains challenging. Method: We propose the Automated Capability Discovery (ACD) framework, which endows large language models with a “scientist” role to autonomously generate open-ended tasks and evaluate capabilities—and failure modes—of target models (including themselves), enabling fully automated, human-free capability mapping. ACD introduces a novel model-driven probing paradigm integrating dual generative–evaluative loops and consistency calibration, validated via human evaluation to establish a reliable automated scoring system. Results: ACD automatically identifies thousands of implicit capabilities and failure patterns across multiple model families. Automated scores exhibit strong agreement with human judgments (Cohen’s κ > 0.85). The implementation, including full experimental logs, is publicly released.
Traditional AI evaluation methods are primarily designed for static model selection and often fail to diagnose root causes of performance degradation in production or guide targeted improvements. This work proposes EvalLoop, a novel methodology that embeds evaluation into a continuous optimization loop. By integrating dimensional metric grouping, failure mode categorization, single-variable controlled experiments, and a human-in-the-loop gating mechanism, EvalLoop enables precise mapping from failure attribution to actionable refinements and supports deployment-aware model selection. Evaluated on a sales intelligence briefing generation task, the approach increased overall accuracy of the best-performing model from 82.6% to 94.6%, improved performance on critical dimensions by over 16 percentage points, and reduced human review effort by 94%.
Rapid AI model evolution has led to ad hoc, non-reproducible model selection in scientific software engineering, severely undermining reproducibility and transparency. To address this, we propose ModelSelect—the first evidence-driven framework that formalizes AI model selection as a multi-criteria decision-making (MCDM) problem, integrating automated metadata harvesting, a structured knowledge graph, and research-context-aware decision modeling. Its key contributions are: (1) establishing the first MCDM modeling paradigm tailored to scientific practice; (2) end-to-end integration of technical metrics and domain semantics, significantly enhancing recommendation interpretability and consistency; and (3) empirical validation across 50 real-world research scenarios, achieving 96.2% coverage, substantially higher rationale alignment than baselines, and superior traceability, cross-scenario robustness, and transparency.
This work addresses the challenge software engineers face in efficiently identifying suitable pre-trained models and datasets for software engineering tasks amid the vast landscape of machine learning assets. To bridge this gap, we propose and implement MLAssetSelection—the first asset selection tool tailored specifically for the software engineering domain. By automatically harvesting relevant assets from platforms such as Hugging Face and integrating multidimensional evaluation metrics, a configurable leaderboard, requirement-based filtering mechanisms, and personalized recommendations, the system enables real-time updates and user customization. Empirical evaluation demonstrates that MLAssetSelection substantially improves the efficiency of model and dataset selection, thereby filling a critical void in domain-specific intelligent asset curation tools.
Traditional AI systems rely on fixed monolithic models, which struggle to dynamically allocate resources, decompose tasks, or update knowledge in response to varying inputs, leading to degraded performance and increased costs. This work proposes the first system-level design methodology for distributed composite AI systems, formulating a design space through workflow topologies and configuration choices and identifying eight core design patterns. The framework jointly optimizes model selection and runtime parameters, enabling task decomposition, multi-model orchestration, and explicit control logic, thereby facilitating a shift from static monolithic architectures toward dynamic, composable, and adaptive ones. Evaluated across three case studies, the approach reduces latency by up to 60% and cost by up to 71%, with only a 2.5–4 percentage point drop in accuracy.
This study addresses the underexplored engineering challenges in existing machine learning evaluation frameworks, where operational issues and their root causes have lacked systematic investigation. To bridge this gap, the work formally establishes evaluation engineering as a distinct research direction within software engineering. Through an empirical analysis of 57 frameworks and a comprehensive categorization of 16,560 reported issues across a newly proposed five-stage workflow model, the study reveals that 41.4% of problems originate in the specification phase, while 61.7% of classified issues stem from missing functionality, inadequate documentation, and insufficient input validation. The findings yield a structured taxonomy of evaluation-related problems and provide empirical evidence to inform the design and improvement of robust evaluation systems.
Current IDEs lack intelligent, end-to-end support for the machine learning (ML) lifecycle, while MLOps platforms remain decoupled from coding environments. To bridge this gap, we propose a novel large language model (LLM)-enhanced intelligent IDE paradigm that deeply integrates LLMs into the development environment. This enables synergistic, closed-loop automation across code-level intelligent programming—such as code generation, debugging, and completion—and full-stack MLOps pipeline orchestration—including data validation, feature store management, data drift detection, retraining triggers, and CI/CD deployment. The system unifies development, experimentation, validation, and monitoring phases, significantly improving engineering efficiency and reproducibility. Empirical evaluation on the UCI Adult and M5 datasets demonstrates a 61% reduction in pipeline configuration time, a 45% improvement in experimental reproducibility, and a 14% increase in data drift detection accuracy.