Score
Designs, implements, and analyzes the procedures and systems that fit machine learning models to data, including defining objectives and loss functions, selecting and implementing optimization algorithms, preparing and batching training data, constructing training pipelines (including distributed or online training), tuning hyperparameters and regularization, and instrumenting evaluation, monitoring, and convergence analysis.
Medical AI deployment is hindered by insufficient production readiness of machine learning (ML) training pipelines. Method: This paper presents a progressive architectural evolution path—monolithic (chaotic) → modular monolithic → microservices—using SPIRA, a voice-based pre-diagnostic system for respiratory insufficiency, as a case study. It systematically introduces continuous training (CT) and a software-quality-attribute-driven MLOps governance framework tailored to healthcare, integrating modular design, microservice decomposition, and engineered CI/CD pipelines. Contribution/Results: The approach significantly improves pipeline maintainability, fault tolerance, and scalability, enabling stable, iterative evolution of SPIRA. It establishes an “agile ML + robust software engineering” co-design paradigm, delivering a reusable methodology and practical benchmark for engineering medical AI in highly regulated environments.
This paper identifies a systemic issue in machine learning: preprocessing hyperparameters—such as missing-value imputation strategies—are frequently overlooked yet substantially bias model evaluation. Current practice often involves informal, post-hoc tuning of preprocessing steps, leading to optimistic performance estimates and irreproducible results. To address this, the authors formally distinguish and empirically analyze the coupling effects between algorithmic and preprocessing hyperparameters. Using a modular supervised learning workflow model, controlled variable experiments, replication of canonical case studies, and bias diagnostics, they quantify the resulting optimistic bias. Key contributions include: (1) establishing preprocessing hyperparameters as equally critical as algorithmic ones; (2) proposing formal modeling principles to eliminate informal preprocessing tuning; and (3) delivering actionable reporting guidelines for ML practitioners, thereby significantly enhancing model credibility and reproducibility.
To address the challenge non-experts face in systematically identifying data requirements for machine learning systems, this paper proposes a goal-oriented modeling approach for data requirement elicitation, introducing the customizable Data Requirement Goal Model (DRGM). DRGM integrates Goal-Question-Metric (GQM)-inspired GRL-based modeling with a dynamic customization mechanism driven by gray and white literature, enabling flexible configuration by task type, KPI metrics, and goal weights. It supports structured assessment of data attribute quality and contextual suitability. As the first dedicated goal model targeting ML-specific data requirements, DRGM bridges a critical methodological gap in involving non-experts in data requirement engineering. Empirical validation across two real-world projects demonstrates high alignment between DRGM-derived requirements and actual needs, significantly improving non-experts’ accuracy, interpretability, and decision-support capability in specifying data requirements and comparing alternative datasets.
This paper addresses the inefficiency and lack of scalability of manual hyperparameter tuning in large-scale machine learning. It systematically surveys hyperparameter optimization (HPO), unifying and classifying five mainstream paradigms: random/low-discrepancy search, bandit-based methods, Bayesian optimization, population-based (evolutionary) algorithms, and gradient-based differentiable optimization. The survey further extends to emerging settings—including online HPO, constrained HPO, and multi-objective HPO. Crucially, the work establishes novel theoretical connections between HPO and meta-learning as well as neural architecture search, yielding a comprehensive knowledge framework that articulates methodological principles, applicability boundaries, and inherent limitations. By clarifying the technical evolution and identifying key open challenges, this study provides a theoretically grounded yet practically actionable foundation for automated machine learning.
To address the inefficiency and poor generalizability of manual hyperparameter tuning—particularly for learning rates—this paper proposes a dynamic online meta-optimization framework that formulates learning rate adaptation as a discounted cumulative regret minimization problem over time. The method employs a gradient-based meta-update mechanism, enabling plug-and-play integration with any first-order optimizer (e.g., SGD, Adam) to achieve decoupled, real-time, adaptive step-size optimization. Key contributions include: (i) the first formalization of meta-optimization as discounted regret minimization; and (ii) a low-complexity variant that preserves theoretical rigor while ensuring computational efficiency and strong generalization. Experiments across diverse tasks demonstrate faster convergence, enhanced robustness to initialization and task heterogeneity, competitive performance against hand-tuned optimal schedulers, and significantly lower computational overhead compared to conventional hyperparameter search methods.
This work proposes the first end-to-end automated artificial intelligence research framework capable of fully automating the development pipeline from algorithmic idea generation to executable machine learning classifiers. The approach integrates structured meta-prompt engineering with large language model–based code generation, augmented by an automated evaluation and iterative optimization mechanism. Experimental results on twenty standard datasets from the Infinity-Bench benchmark demonstrate that multiple novel classifiers autonomously generated by the framework significantly outperform baseline methods implemented in scikit-learn. This study thus achieves, for the first time, complete automation of the entire workflow—from initial algorithmic conception to deployable, runnable code—marking a significant step toward self-driving AI research systems.
This study addresses the prevailing gap in AI education, which emphasizes model development while neglecting system engineering practices, leaving students ill-equipped to handle real-world challenges such as architectural design, deployment, and monitoring. To bridge this gap, the authors implemented a master’s-level course in which students built a movie recommendation system under realistic constraints, with a focus on integrating AI components into robust software systems, adopting data-driven machine learning practices, and cultivating systems-level thinking. Using a mixed-methods approach—combining analysis of student project artifacts with survey data—the research evaluates learners’ performance in architectural decision-making, integration of heterogeneous models, and adaptation to evolving requirements. Findings reveal common difficulties students encounter in AI system engineering and demonstrate the course’s effectiveness in addressing critical deficiencies in AI engineering education and enhancing systems-aware competencies.
This study addresses the prevalent ad hoc and non-standardized practices in model integration and deployment within MLOps projects, which often stem from a lack of systematic architectural guidance. To bridge this gap, the authors conduct a gray literature review of 103 online sources and apply thematic analysis to derive, for the first time, 25 architecturally significant best practices. These practices are systematically categorized into five thematic groups, with explicit articulation of each practice’s impact on overall system architecture. The resulting framework offers a structured, actionable set of guidelines for MLOps model integration and deployment, providing both researchers and engineering teams with a coherent theoretical foundation and practical reference for designing robust, scalable machine learning systems.
This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.
To address model performance degradation caused by data distribution drift and the reliance of existing MLOps retraining pipelines on manual intervention, this paper proposes an automated, adaptive neural network retraining framework. Methodologically, it introduces a novel multi-criteria joint drift detection mechanism—integrating statistical metrics including the Kolmogorov–Smirnov test, Population Stability Index (PSI), and Classifier-Driven (CD) drift detection—combined with online monitoring and lightweight scheduling to dynamically trigger end-to-end retraining upon significant drift. The framework is implemented using a cloud-native architecture for scalable and efficient deployment. Evaluated on multiple benchmark datasets, the proposed approach improves classification accuracy by 12.3%–18.7%, reduces inference latency by 41%, and cuts computational resource consumption by 53%, compared to conventional periodic or single-threshold retraining strategies. These gains significantly enhance model freshness and operational cost-efficiency.