Score
Designs, builds, and operates end-to-end machine learning systems, including data ingestion and preprocessing pipelines, model training and evaluation workflows, deployment and serving infrastructure, and monitoring for performance, reliability, and reproducibility. Implements scalable, maintainable engineering practices (versioning, CI/CD, testing, resource management) to integrate trained models into production services and data-driven applications.
Real-world deployment of AI systems faces dual challenges: heterogeneous data floods and stringent low-latency requirements—straining conventional software architectures to their limits. To address this, we conduct the first systematic analysis of 217 production-deployed ML systems from a Data-Oriented Architecture (DOA) perspective, uncovering implicit DOA design patterns—particularly in loose coupling, decentralization, and data-driven orchestration. Leveraging systematic literature review, architectural pattern extraction, and requirement-to-design mapping analysis, we synthesize the first empirically grounded DOA practice guide and open-challenge taxonomy tailored to real-world ML deployments. Furthermore, we propose actionable, reusable recommendations for ML system deployment and introduce a novel architecture evaluation framework. Collectively, these contributions bridge the critical knowledge gap between DOA theory and industrial engineering practice.
Medical AI deployment is hindered by insufficient production readiness of machine learning (ML) training pipelines. Method: This paper presents a progressive architectural evolution path—monolithic (chaotic) → modular monolithic → microservices—using SPIRA, a voice-based pre-diagnostic system for respiratory insufficiency, as a case study. It systematically introduces continuous training (CT) and a software-quality-attribute-driven MLOps governance framework tailored to healthcare, integrating modular design, microservice decomposition, and engineered CI/CD pipelines. Contribution/Results: The approach significantly improves pipeline maintainability, fault tolerance, and scalability, enabling stable, iterative evolution of SPIRA. It establishes an “agile ML + robust software engineering” co-design paradigm, delivering a reusable methodology and practical benchmark for engineering medical AI in highly regulated environments.
This study addresses unique challenges hindering continuous integration (CI) adoption in machine learning (ML) projects—namely, long build times, low test coverage, non-deterministic behavior, and strong data dependencies. Using a mixed-methods approach, we conducted a survey of 155 ML practitioners, performed in-depth interviews with thematic coding, and analyzed empirical data from 47 industrial ML projects. Our analysis identified eight key distinctions between traditional and ML-specific CI and five recurring challenges. We propose novel, ML-tailored CI practices: automated tracking of model performance metrics, dynamic test prioritization, and interdisciplinary collaboration–driven mechanisms to strengthen testing culture. Finally, we synthesize these insights into a practical, empirically grounded ML-CI best-practices guide—the first systematic, evidence-based framework for building efficient and robust CI pipelines in ML development.
This study addresses the inherent tension between scalability and maintainability in machine learning (ML) systems—a critical challenge impeding robust, production-grade deployment. Method: We conduct a systematic literature review (SLR) grounded in 124 high-quality publications, developing the first six-dimensional analytical framework spanning data engineering, model engineering, and system deployment. Contribution/Results: The work identifies 41 categories of maintainability issues and 13 categories of scalability issues, uncovering their stage-crossing trade-offs and synergies. It introduces the first taxonomy of scalability–maintainability challenges in ML systems, accompanied by a problem distribution map and an evidence-based repository quantifying solution effectiveness. Collectively, these findings deliver empirically grounded, cross-stage design principles and actionable optimization pathways for industrial ML system development.
This study addresses the challenges of integrating machine learning (ML) models into software systems—namely, poor integration practices, low reusability, and unclear architectural boundaries. It presents the first large-scale empirical investigation across 2,928 open-source ML-enabled systems. Leveraging GitHub code mining, static analysis, topic modeling, and architectural pattern identification, the work systematically characterizes ML integration topologies, code/model reuse practices, and maintenance bottlenecks. Key contributions include: (1) the first comprehensive classification framework and architectural pattern atlas for ML-enabled systems; (2) identification of seven prevalent integration topologies and four model reuse patterns; and (3) uncovering critical interdisciplinary collaboration barriers in ML-software co-development. The findings bridge the methodological gap between data science and software engineering at the model embedding stage, providing industry-practical architectural guidelines that significantly enhance the maintainability and reusability of ML systems.
This study presents the first empirical investigation into the evolution of CI/CD configurations in machine learning (ML) projects. Addressing the lack of understanding regarding how CI/CD configurations co-evolve with ML components, the authors analyze 508 open-source ML projects, 343 manually annotated commits, and 15,634 automated CI/CD commits. They propose a novel 14-category taxonomy capturing synergistic changes between CI/CD and ML components, develop a dedicated clustering tool to identify recurrent evolutionary patterns, and establish an empirically grounded model linking developer experience to CI/CD configuration modification behavior. Results show that 61.8% of CI/CD-related commits involve build strategy modifications; common anti-patterns—including dependency hardcoding and missing test frameworks—are identified; and senior developers modify CI/CD configurations more frequently and effectively than juniors, confirming the critical role of experience in CI/CD maintenance.
This work addresses customs clearance delays in global trade caused by ambiguous product descriptions and frequent updates to Harmonized System (HS) codes. To tackle this challenge, the authors propose a serverless MLOps framework that leverages event-driven pipelines and managed services to enable end-to-end, model-agnostic machine learning lifecycle management. The architecture supports automatic scaling, reproducible training, auditable deployment, and automated A/B testing, ensuring secure and seamless model transitions. By integrating custom text embeddings with models such as Text-CNN, the system achieves 98% accuracy on real-world HS code prediction tasks, meeting stringent service-level agreement (SLA) requirements. This approach significantly reduces long-term operational costs and establishes an efficient, cost-effective, and reproducible deployment paradigm for industrial-scale machine learning systems.
This study addresses the prevailing gap in AI education, which emphasizes model development while neglecting system engineering practices, leaving students ill-equipped to handle real-world challenges such as architectural design, deployment, and monitoring. To bridge this gap, the authors implemented a master’s-level course in which students built a movie recommendation system under realistic constraints, with a focus on integrating AI components into robust software systems, adopting data-driven machine learning practices, and cultivating systems-level thinking. Using a mixed-methods approach—combining analysis of student project artifacts with survey data—the research evaluates learners’ performance in architectural decision-making, integration of heterogeneous models, and adaptation to evolving requirements. Findings reveal common difficulties students encounter in AI system engineering and demonstrate the course’s effectiveness in addressing critical deficiencies in AI engineering education and enhancing systems-aware competencies.
This study addresses the prevalent ad hoc and non-standardized practices in model integration and deployment within MLOps projects, which often stem from a lack of systematic architectural guidance. To bridge this gap, the authors conduct a gray literature review of 103 online sources and apply thematic analysis to derive, for the first time, 25 architecturally significant best practices. These practices are systematically categorized into five thematic groups, with explicit articulation of each practice’s impact on overall system architecture. The resulting framework offers a structured, actionable set of guidelines for MLOps model integration and deployment, providing both researchers and engineering teams with a coherent theoretical foundation and practical reference for designing robust, scalable machine learning systems.
This work addresses the challenges of SLO violations and resource inefficiency in machine learning model serving caused by inadequate capacity planning. To this end, the authors propose an adaptive, feedback-driven load testing framework that formalizes the ML serving load testing process for the first time. The framework incorporates real-traffic-based workload calibration and a warm-up mechanism, combined with adaptive search, performance signal feedback control, convergence detection, and GPU monitoring to efficiently estimate the maximum sustainable throughput under SLO constraints. Evaluation across 14 industrial cases demonstrates that the approach reduces capacity estimation error from approximately 30% to 2–6%, with the warm-up mechanism improving accuracy by 22.2%. This significantly mitigates deployment incidents and enhances GPU resource utilization efficiency.