Score
Designs, builds, and operates end-to-end machine learning production systems and workflows, including automated training pipelines, model packaging and deployment, CI/CD for models, data and model versioning, orchestration, and infrastructure for scalable inference. Uses tooling and processes to monitor model performance and reliability, implement observability and alerts, manage reproducibility and lineage, and automate retraining and lifecycle governance.
To address the lack of a unified knowledge framework in MLOps, this paper conducts a multi-source literature review (MLR), systematically synthesizing 150 academic publications and 48 grey literature sources to overcome single-perspective limitations. Through thematic coding and cross-source evidence triangulation, it establishes the first comprehensive MLOps conceptual model and practice map spanning the full ML lifecycle and integrating consensus from both industry and academia. Key contributions include: (1) a widely adopted, rigorous definition of MLOps; (2) distillation of 12 core MLOps practices; and (3) identification of seven recurrent implementation challenges alongside empirically grounded mitigation strategies. The resulting knowledge base is modular, reusable, and rigorously validated—serving as a foundational reference for MLOps standardization, tooling development, and empirical research.
Medical AI deployment is hindered by insufficient production readiness of machine learning (ML) training pipelines. Method: This paper presents a progressive architectural evolution path—monolithic (chaotic) → modular monolithic → microservices—using SPIRA, a voice-based pre-diagnostic system for respiratory insufficiency, as a case study. It systematically introduces continuous training (CT) and a software-quality-attribute-driven MLOps governance framework tailored to healthcare, integrating modular design, microservice decomposition, and engineered CI/CD pipelines. Contribution/Results: The approach significantly improves pipeline maintainability, fault tolerance, and scalability, enabling stable, iterative evolution of SPIRA. It establishes an “agile ML + robust software engineering” co-design paradigm, delivering a reusable methodology and practical benchmark for engineering medical AI in highly regulated environments.
This study presents the first empirical investigation into the evolution of CI/CD configurations in machine learning (ML) projects. Addressing the lack of understanding regarding how CI/CD configurations co-evolve with ML components, the authors analyze 508 open-source ML projects, 343 manually annotated commits, and 15,634 automated CI/CD commits. They propose a novel 14-category taxonomy capturing synergistic changes between CI/CD and ML components, develop a dedicated clustering tool to identify recurrent evolutionary patterns, and establish an empirically grounded model linking developer experience to CI/CD configuration modification behavior. Results show that 61.8% of CI/CD-related commits involve build strategy modifications; common anti-patterns—including dependency hardcoding and missing test frameworks—are identified; and senior developers modify CI/CD configurations more frequently and effectively than juniors, confirming the critical role of experience in CI/CD maintenance.
This study addresses unique challenges hindering continuous integration (CI) adoption in machine learning (ML) projects—namely, long build times, low test coverage, non-deterministic behavior, and strong data dependencies. Using a mixed-methods approach, we conducted a survey of 155 ML practitioners, performed in-depth interviews with thematic coding, and analyzed empirical data from 47 industrial ML projects. Our analysis identified eight key distinctions between traditional and ML-specific CI and five recurring challenges. We propose novel, ML-tailored CI practices: automated tracking of model performance metrics, dynamic test prioritization, and interdisciplinary collaboration–driven mechanisms to strengthen testing culture. Finally, we synthesize these insights into a practical, empirically grounded ML-CI best-practices guide—the first systematic, evidence-based framework for building efficient and robust CI pipelines in ML development.
Current IDEs lack intelligent, end-to-end support for the machine learning (ML) lifecycle, while MLOps platforms remain decoupled from coding environments. To bridge this gap, we propose a novel large language model (LLM)-enhanced intelligent IDE paradigm that deeply integrates LLMs into the development environment. This enables synergistic, closed-loop automation across code-level intelligent programming—such as code generation, debugging, and completion—and full-stack MLOps pipeline orchestration—including data validation, feature store management, data drift detection, retraining triggers, and CI/CD deployment. The system unifies development, experimentation, validation, and monitoring phases, significantly improving engineering efficiency and reproducibility. Empirical evaluation on the UCI Adult and M5 datasets demonstrates a 61% reduction in pipeline configuration time, a 45% improvement in experimental reproducibility, and a 14% increase in data drift detection accuracy.
To address inefficiencies in model incremental updates, unfair policy evaluation, and high retraining costs under continual data growth, this paper proposes an end-to-end adaptive machine learning platform. Methodologically: (1) it introduces a declarative domain-specific language (DSL) to uniformly model data selection strategies (e.g., coreset, uncertainty sampling) and trigger policies (e.g., drift-aware scheduling); (2) it establishes the first composite model evaluation framework enabling fair, cross-policy comparison; and (3) it implements a co-optimization mechanism integrating sample-level fine-grained data selection with high-throughput training. Contributions include an open-source, extensible system architecture, a standardized benchmark ecosystem, and abstracted ML pipeline interfaces. Experiments demonstrate significant improvements in training throughput and substantial reductions in retraining overhead—while preserving model accuracy—and enable reproducible analysis across diverse strategy combinations.
This work addresses the heavy reliance on expert knowledge in designing and debugging scientific workflows, a challenge exacerbated by existing large language model approaches that directly generate code without ensuring transparency, reproducibility, or seamless system integration. To overcome these limitations, we propose an AI-assisted scientific workflow management framework that decouples user intent from implementation through a structured specification phase, enabling specification-driven workflow generation and validation. We further introduce a multi-layer debugging agent powered by large language models to automate error diagnosis and correction. By deeply integrating with the Pegasus workflow system via the Model Context Protocol (MCP), our approach supports end-to-end workflow lifecycle management. Empirical evaluation demonstrates successful generation and execution of federated learning medical imaging workflows comprising thousands of tasks, substantially reducing debugging effort and empowering non-expert users to construct complex workflows adhering to expert-level design patterns.
This work addresses the limitations of existing CI/CD workflow analyses, which often focus narrowly on stage identification and struggle to assess reliability, maintainability, and optimization priorities. To overcome this, we propose a large language model–based CI/CD analysis pipeline that integrates repository context enhancement, anti-pattern detection, stage mining, and actionable recommendation generation. Our approach uniquely combines diagnostic reasoning, context awareness, and human-in-the-loop review to deliver observability tailored to cybersecurity engineering. Leveraging few-shot prompting, YAML parsing, and statistical tests (chi-square and Cramér’s V), the method identifies 434,769 anti-patterns across 75,201 workflows and generates an average of 8.25 syntactically valid optimization suggestions per repository, achieving a 96.1% compliance rate with YAML syntax standards.
This work addresses the growing complexity of CI/CD pipelines and the lack of structured analysis capabilities in existing tools for understanding their behavior, failures, and version evolution. The authors propose an innovative approach that uniquely integrates digital twin technology with BPMN-based modeling in DevOps contexts. By automatically parsing raw CI configurations and execution logs, the method constructs structured, high-level process models that enable pipeline visualization, failure traceability, and cross-version comparison. Evaluated across multiple open-source projects, the approach demonstrates effectiveness in monitoring, evolutionary analysis, and fault diagnosis, offering a modular and extensible foundational framework for the analysis and optimization of CI/CD pipelines.
This study addresses the prevalent ad hoc and non-standardized practices in model integration and deployment within MLOps projects, which often stem from a lack of systematic architectural guidance. To bridge this gap, the authors conduct a gray literature review of 103 online sources and apply thematic analysis to derive, for the first time, 25 architecturally significant best practices. These practices are systematically categorized into five thematic groups, with explicit articulation of each practice’s impact on overall system architecture. The resulting framework offers a structured, actionable set of guidelines for MLOps model integration and deployment, providing both researchers and engineering teams with a coherent theoretical foundation and practical reference for designing robust, scalable machine learning systems.