Score
Design and implement training procedures, algorithms, and evaluation pipelines that first fit a model using pre-collected (offline) data and then continue improving it through live (online) updates; this includes the two-step scheduling, mechanisms to leverage offline pretraining, techniques to mitigate distribution shift between offline and online data, and methods to ensure stable, sample-efficient adaptation during deployment.
This paper addresses five core challenges in dynamic data environments—data drift, concept drift, catastrophic forgetting, skewed learning, and network adaptability. Method: It systematically surveys over 120 state-of-the-art evolutionary machine learning (EML) works, integrating online learning, incremental learning, continual learning, meta-learning, dynamic pruning, and ensemble distillation to establish a multi-paradigm evaluation framework covering supervised, unsupervised, and semi-supervised settings. Contribution/Results: The work introduces the first unified analytical framework for EML, clarifies challenge taxonomies, uncovers synergistic mechanisms among adaptive neural architectures, meta-learning, and ensemble strategies, and identifies critical gaps in robustness, ethics, and scalability. It delivers a comprehensive EML methodology landscape, a curated collection of mainstream benchmarks and evaluation metrics, and system design principles tailored for industrial deployment—providing both theoretical foundations and practical guidance for building dynamic AI systems.
To address degraded model adaptability in online multistep time-series forecasting—caused by data distribution drift and delayed ground-truth feedback—this paper proposes ADAPT-Z. Methodologically, ADAPT-Z abandons conventional parameter fine-tuning and instead models the dynamics of latent factors. It introduces an adapter module that fuses current features with historical gradient information within a learned Z-space, enabling persistent tracking and incremental self-adaptation of feature representations. This design mitigates gradient mismatch induced by label delay and enhances robustness to non-stationary data. Empirical evaluation across multiple benchmark datasets demonstrates that ADAPT-Z significantly outperforms static baselines and state-of-the-art online learning methods, achieving superior generalization and sustained adaptive capability under streaming conditions.
To address model performance degradation caused by data distribution drift and the reliance of existing MLOps retraining pipelines on manual intervention, this paper proposes an automated, adaptive neural network retraining framework. Methodologically, it introduces a novel multi-criteria joint drift detection mechanism—integrating statistical metrics including the Kolmogorov–Smirnov test, Population Stability Index (PSI), and Classifier-Driven (CD) drift detection—combined with online monitoring and lightweight scheduling to dynamically trigger end-to-end retraining upon significant drift. The framework is implemented using a cloud-native architecture for scalable and efficient deployment. Evaluated on multiple benchmark datasets, the proposed approach improves classification accuracy by 12.3%–18.7%, reduces inference latency by 41%, and cuts computational resource consumption by 53%, compared to conventional periodic or single-threshold retraining strategies. These gains significantly enhance model freshness and operational cost-efficiency.
Online fine-tuning of offline pre-trained RL models typically requires continuous access to large-scale offline datasets, incurring high computational overhead, slow convergence, and risks of Q-function divergence and catastrophic forgetting due to distributional shift. Method: We theoretically establish, for the first time, that offline data are unnecessary during online fine-tuning, and propose Warm-start RL (WSRL)—a novel paradigm that initiates online adaptation using only a small number of rollouts generated by the pre-trained policy. WSRL integrates policy warmup, distribution-matching analysis, and an offline-to-online policy bridging mechanism, eliminating the need to store or revisit any offline data. Contribution/Results: Evaluated across multiple standard benchmarks, WSRL consistently outperforms state-of-the-art methods—both those retaining and discarding offline data—in final performance and sample efficiency. It accelerates convergence by 30–50%, achieves higher asymptotic returns, and reduces training cost by an order of magnitude.
Medical AI deployment is hindered by insufficient production readiness of machine learning (ML) training pipelines. Method: This paper presents a progressive architectural evolution path—monolithic (chaotic) → modular monolithic → microservices—using SPIRA, a voice-based pre-diagnostic system for respiratory insufficiency, as a case study. It systematically introduces continuous training (CT) and a software-quality-attribute-driven MLOps governance framework tailored to healthcare, integrating modular design, microservice decomposition, and engineered CI/CD pipelines. Contribution/Results: The approach significantly improves pipeline maintainability, fault tolerance, and scalability, enabling stable, iterative evolution of SPIRA. It establishes an “agile ML + robust software engineering” co-design paradigm, delivering a reusable methodology and practical benchmark for engineering medical AI in highly regulated environments.
Existing methods for online nonparametric estimation on streaming data lack efficient, adaptive hyperparameter selection mechanisms. Method: We propose Weighted Rolling Validation (WRV), a low-overhead, fully online model selection framework that generalizes leave-one-out cross-validation to the streaming setting via temporal weighting of historical validation samples. Grounded in statistical stability assumptions, WRV dynamically assigns time-decaying weights without requiring additional storage or retraining, and is compatible with stochastic gradient–based nonparametric estimators. Contribution/Results: We establish theoretical guarantees showing WRV achieves adaptive convergence rates. Empirically, WRV exhibits high sensitivity to subtle performance differences among candidate estimators, incurs negligible computational overhead, and significantly improves prediction accuracy and robustness. To our knowledge, WRV is the first lightweight, theoretically grounded hyperparameter adaptation mechanism for online nonparametric learning.
This work addresses the fragmented landscape of post-training adaptation techniques, which suffer from inconsistent terminology and a lack of unified comparative or governance frameworks. To resolve this, the paper introduces the first six-dimensional taxonomy—spanning mechanism, objective, data requirements, persistence, structural scope, and model type—that systematically integrates mainstream approaches such as fine-tuning, retrieval augmentation, prompt engineering, model editing, and machine unlearning. This framework clarifies conceptual boundaries and reveals evolutionary and compositional relationships among methods. Beyond standardizing terminology, it enables standardized technical documentation, model change tracking, and AI governance analysis. The study further identifies critical challenges, including evaluation rigor, reproducibility, continual adaptation, multimodal alignment, and governance-aware workflows.
This work addresses the lack of a general, auditable dynamic control mechanism in existing training systems, which typically rely on framework-specific code. The authors propose the first cross-framework, open-source control plane that exposes training interfaces through a unified protocol, integrating declarative configuration, request validation, and secure control-point scheduling within the Aim workspace to enable metric monitoring, real-time intervention, and operational traceability. The system supports safe human and automated controller interventions during training while fully logging all operational trajectories. Experiments across five NLP and reinforcement learning tasks demonstrate its effectiveness, and the open-source implementation provides a foundation for reproducible human-in-the-loop training.
In industrial MLOps, machine learning models often degrade due to data drift yet lack systematic mechanisms for timely updates. Addressing this challenge, this work proposes and systematically evaluates three transfer learning strategies—Ensemble Transfer Learning (ETL), All-Layer Transfer Learning (ALTL), and Last-Layer Transfer Learning (LLTL)—for updating degraded feedforward neural networks under varying data batch sizes. Experimental results demonstrate that ETL achieves the highest prediction accuracy in small-batch scenarios (e.g., 5-day intervals), whereas ALTL performs better in larger-batch settings (e.g., 8-day intervals). This study provides empirical evidence and practical guidance for selecting efficient, adaptive model update strategies in real-world industrial environments, thereby enhancing model robustness and longevity amid evolving data distributions.