Score
Designs and builds automated, end-to-end systems that package, test, validate, deploy, and monitor machine learning models and data workflows into production, including CI/CD pipelines, orchestration, and integration with infrastructure and serving layers. Implements tooling, frameworks, automation, and monitoring/observability to support deployment, scaling, rollback, versioning, and operational best practices for repeatable, production-safe ML operations.
To address the lack of a unified knowledge framework in MLOps, this paper conducts a multi-source literature review (MLR), systematically synthesizing 150 academic publications and 48 grey literature sources to overcome single-perspective limitations. Through thematic coding and cross-source evidence triangulation, it establishes the first comprehensive MLOps conceptual model and practice map spanning the full ML lifecycle and integrating consensus from both industry and academia. Key contributions include: (1) a widely adopted, rigorous definition of MLOps; (2) distillation of 12 core MLOps practices; and (3) identification of seven recurrent implementation challenges alongside empirically grounded mitigation strategies. The resulting knowledge base is modular, reusable, and rigorously validated—serving as a foundational reference for MLOps standardization, tooling development, and empirical research.
Real-world deployment of AI systems faces dual challenges: heterogeneous data floods and stringent low-latency requirements—straining conventional software architectures to their limits. To address this, we conduct the first systematic analysis of 217 production-deployed ML systems from a Data-Oriented Architecture (DOA) perspective, uncovering implicit DOA design patterns—particularly in loose coupling, decentralization, and data-driven orchestration. Leveraging systematic literature review, architectural pattern extraction, and requirement-to-design mapping analysis, we synthesize the first empirically grounded DOA practice guide and open-challenge taxonomy tailored to real-world ML deployments. Furthermore, we propose actionable, reusable recommendations for ML system deployment and introduce a novel architecture evaluation framework. Collectively, these contributions bridge the critical knowledge gap between DOA theory and industrial engineering practice.
This study presents the first empirical investigation into the evolution of CI/CD configurations in machine learning (ML) projects. Addressing the lack of understanding regarding how CI/CD configurations co-evolve with ML components, the authors analyze 508 open-source ML projects, 343 manually annotated commits, and 15,634 automated CI/CD commits. They propose a novel 14-category taxonomy capturing synergistic changes between CI/CD and ML components, develop a dedicated clustering tool to identify recurrent evolutionary patterns, and establish an empirically grounded model linking developer experience to CI/CD configuration modification behavior. Results show that 61.8% of CI/CD-related commits involve build strategy modifications; common anti-patterns—including dependency hardcoding and missing test frameworks—are identified; and senior developers modify CI/CD configurations more frequently and effectively than juniors, confirming the critical role of experience in CI/CD maintenance.
Medical AI deployment is hindered by insufficient production readiness of machine learning (ML) training pipelines. Method: This paper presents a progressive architectural evolution path—monolithic (chaotic) → modular monolithic → microservices—using SPIRA, a voice-based pre-diagnostic system for respiratory insufficiency, as a case study. It systematically introduces continuous training (CT) and a software-quality-attribute-driven MLOps governance framework tailored to healthcare, integrating modular design, microservice decomposition, and engineered CI/CD pipelines. Contribution/Results: The approach significantly improves pipeline maintainability, fault tolerance, and scalability, enabling stable, iterative evolution of SPIRA. It establishes an “agile ML + robust software engineering” co-design paradigm, delivering a reusable methodology and practical benchmark for engineering medical AI in highly regulated environments.
This study addresses unique challenges hindering continuous integration (CI) adoption in machine learning (ML) projects—namely, long build times, low test coverage, non-deterministic behavior, and strong data dependencies. Using a mixed-methods approach, we conducted a survey of 155 ML practitioners, performed in-depth interviews with thematic coding, and analyzed empirical data from 47 industrial ML projects. Our analysis identified eight key distinctions between traditional and ML-specific CI and five recurring challenges. We propose novel, ML-tailored CI practices: automated tracking of model performance metrics, dynamic test prioritization, and interdisciplinary collaboration–driven mechanisms to strengthen testing culture. Finally, we synthesize these insights into a practical, empirically grounded ML-CI best-practices guide—the first systematic, evidence-based framework for building efficient and robust CI pipelines in ML development.
Small- and medium-sized enterprises (SMEs) face prohibitive costs, high production downtime risks, and operational complexity when integrating machine learning (ML) into legacy industrial systems. Method: This paper proposes a human-in-the-loop interactive ML framework that decouples the ML model lifecycle from the production environment via an API-based middleware layer. It employs a lightweight model-serving architecture and a browser-based interactive interface, enabling zero-hardware-upgrade deployment, zero-downtime integration, and remote real-time parameter tuning. Contribution/Results: The framework is the first to support dynamic model maintenance and online collaborative decision-making by domain experts—without modifying existing systems. Experimental evaluation demonstrates substantial reductions in ML adoption barriers and implementation costs, alongside measurable improvements in manufacturing quality and safety. The solution exhibits strong scalability and engineering practicality for industrial deployment.
Current IDEs lack intelligent, end-to-end support for the machine learning (ML) lifecycle, while MLOps platforms remain decoupled from coding environments. To bridge this gap, we propose a novel large language model (LLM)-enhanced intelligent IDE paradigm that deeply integrates LLMs into the development environment. This enables synergistic, closed-loop automation across code-level intelligent programming—such as code generation, debugging, and completion—and full-stack MLOps pipeline orchestration—including data validation, feature store management, data drift detection, retraining triggers, and CI/CD deployment. The system unifies development, experimentation, validation, and monitoring phases, significantly improving engineering efficiency and reproducibility. Empirical evaluation on the UCI Adult and M5 datasets demonstrates a 61% reduction in pipeline configuration time, a 45% improvement in experimental reproducibility, and a 14% increase in data drift detection accuracy.
This study addresses the widespread neglect of licensing terms and regulatory compliance in the deployment of machine learning models within open-source software, particularly in safety-critical contexts where associated risks are pronounced. The authors present the first systematic investigation of ML usage across 173 open-source projects on GitHub spanning 16 application domains. Through code inspection and contextual analysis, they evaluate each model’s role in decision-making, the presence of risk-mitigation strategies, and adherence to licensing requirements. The findings reveal that certain projects employ ML for high-stakes decisions without complying with applicable license conditions and often lack essential post-processing safeguards. This work uncovers critical compliance blind spots in the open-source ecosystem and provides an empirical foundation for developing compliance guidelines and automated detection tools.
This work addresses the growing complexity of CI/CD pipelines and the lack of structured analysis capabilities in existing tools for understanding their behavior, failures, and version evolution. The authors propose an innovative approach that uniquely integrates digital twin technology with BPMN-based modeling in DevOps contexts. By automatically parsing raw CI configurations and execution logs, the method constructs structured, high-level process models that enable pipeline visualization, failure traceability, and cross-version comparison. Evaluated across multiple open-source projects, the approach demonstrates effectiveness in monitoring, evolutionary analysis, and fault diagnosis, offering a modular and extensible foundational framework for the analysis and optimization of CI/CD pipelines.
This study addresses the prevalent ad hoc and non-standardized practices in model integration and deployment within MLOps projects, which often stem from a lack of systematic architectural guidance. To bridge this gap, the authors conduct a gray literature review of 103 online sources and apply thematic analysis to derive, for the first time, 25 architecturally significant best practices. These practices are systematically categorized into five thematic groups, with explicit articulation of each practice’s impact on overall system architecture. The resulting framework offers a structured, actionable set of guidelines for MLOps model integration and deployment, providing both researchers and engineering teams with a coherent theoretical foundation and practical reference for designing robust, scalable machine learning systems.
This work addresses customs clearance delays in global trade caused by ambiguous product descriptions and frequent updates to Harmonized System (HS) codes. To tackle this challenge, the authors propose a serverless MLOps framework that leverages event-driven pipelines and managed services to enable end-to-end, model-agnostic machine learning lifecycle management. The architecture supports automatic scaling, reproducible training, auditable deployment, and automated A/B testing, ensuring secure and seamless model transitions. By integrating custom text embeddings with models such as Text-CNN, the system achieves 98% accuracy on real-world HS code prediction tasks, meeting stringent service-level agreement (SLA) requirements. This approach significantly reduces long-term operational costs and establishes an efficient, cost-effective, and reproducible deployment paradigm for industrial-scale machine learning systems.