ml engineering

Designs, builds, and operates end-to-end machine learning systems, including data ingestion and preprocessing pipelines, model training and evaluation workflows, deployment and serving infrastructure, and monitoring for performance, reliability, and reproducibility. Implements scalable, maintainable engineering practices (versioning, CI/CD, testing, resource management) to integrate trained models into production services and data-driven applications.

mlengineering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.51
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$224K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Is Your Training Pipeline Production-Ready? A Case Study in the Healthcare Domain

Jun 07, 2025
DL
Daniel Lawand
🏛️ University of São Paulo | Tilburg University | Technical University of Eindhoven

Medical AI deployment is hindered by insufficient production readiness of machine learning (ML) training pipelines. Method: This paper presents a progressive architectural evolution path—monolithic (chaotic) → modular monolithic → microservices—using SPIRA, a voice-based pre-diagnostic system for respiratory insufficiency, as a case study. It systematically introduces continuous training (CT) and a software-quality-attribute-driven MLOps governance framework tailored to healthcare, integrating modular design, microservice decomposition, and engineered CI/CD pipelines. Contribution/Results: The approach significantly improves pipeline maintainability, fault tolerance, and scalability, enabling stable, iterative evolution of SPIRA. It establishes an “agile ML + robust software engineering” co-design paradigm, delivering a reusable methodology and practical benchmark for engineering medical AI in highly regulated environments.

Ensuring ML training pipelines are production-ready in healthcareEvolving architecture for better maintainability and robustnessImproving software quality in MLES for respiratory pre-diagnosis

Continuous Integration Practices in Machine Learning Projects: The Practitioners` Perspective

Feb 24, 2025
JH
João Helis Bernardo
🏛️ Federal University of Rio Grande do Norte | University of Otago | Huawei

This study addresses unique challenges hindering continuous integration (CI) adoption in machine learning (ML) projects—namely, long build times, low test coverage, non-deterministic behavior, and strong data dependencies. Using a mixed-methods approach, we conducted a survey of 155 ML practitioners, performed in-depth interviews with thematic coding, and analyzed empirical data from 47 industrial ML projects. Our analysis identified eight key distinctions between traditional and ML-specific CI and five recurring challenges. We propose novel, ML-tailored CI practices: automated tracking of model performance metrics, dynamic test prioritization, and interdisciplinary collaboration–driven mechanisms to strengthen testing culture. Finally, we synthesize these insights into a practical, empirically grounded ML-CI best-practices guide—the first systematic, evidence-based framework for building efficient and robust CI pipelines in ML development.

Challenges in applying CI to ML projectsProposing ML-specific CI practicesUnique ML project characteristics and CI

This study addresses the inherent tension between scalability and maintainability in machine learning (ML) systems—a critical challenge impeding robust, production-grade deployment. Method: We conduct a systematic literature review (SLR) grounded in 124 high-quality publications, developing the first six-dimensional analytical framework spanning data engineering, model engineering, and system deployment. Contribution/Results: The work identifies 41 categories of maintainability issues and 13 categories of scalability issues, uncovering their stage-crossing trade-offs and synergies. It introduces the first taxonomy of scalability–maintainability challenges in ML systems, accompanied by a problem distribution map and an evidence-based repository quantifying solution effectiveness. Collectively, these findings deliver empirically grounded, cross-stage design principles and actionable optimization pathways for industrial ML system development.

Analyzes solutions across ML lifecycle from data to deploymentExplores interdependencies between scalability and maintainability in MLIdentifies scalability and maintainability challenges in ML systems

A Large-Scale Study of Model Integration in ML-Enabled Software Systems

Aug 12, 2024
YS
Yorick Sens
🏛️ Ruhr University Bochum | Chalmers University of Gothenburg

This study addresses the challenges of integrating machine learning (ML) models into software systems—namely, poor integration practices, low reusability, and unclear architectural boundaries. It presents the first large-scale empirical investigation across 2,928 open-source ML-enabled systems. Leveraging GitHub code mining, static analysis, topic modeling, and architectural pattern identification, the work systematically characterizes ML integration topologies, code/model reuse practices, and maintenance bottlenecks. Key contributions include: (1) the first comprehensive classification framework and architectural pattern atlas for ML-enabled systems; (2) identification of seven prevalent integration topologies and four model reuse patterns; and (3) uncovering critical interdisciplinary collaboration barriers in ML-software co-development. The findings bridge the methodological gap between data science and software engineering at the model embedding stage, providing industry-practical architectural guidelines that significantly enhance the maintainability and reusability of ML systems.

ML and code reuse practicesML model integration challengesML-enabled system characteristics

Empirical Analysis on CI/CD Pipeline Evolution in Machine Learning Projects

Mar 18, 2024
AH
Alaa Houerbi
🏛️ University of Michigan- Dearborn

This study presents the first empirical investigation into the evolution of CI/CD configurations in machine learning (ML) projects. Addressing the lack of understanding regarding how CI/CD configurations co-evolve with ML components, the authors analyze 508 open-source ML projects, 343 manually annotated commits, and 15,634 automated CI/CD commits. They propose a novel 14-category taxonomy capturing synergistic changes between CI/CD and ML components, develop a dedicated clustering tool to identify recurrent evolutionary patterns, and establish an empirically grounded model linking developer experience to CI/CD configuration modification behavior. Results show that 61.8% of CI/CD-related commits involve build strategy modifications; common anti-patterns—including dependency hardcoding and missing test frameworks—are identified; and senior developers modify CI/CD configurations more frequently and effectively than juniors, confirming the critical role of experience in CI/CD maintenance.

Analyzes CI/CD evolution in ML projectsDevelops clustering tool for CI/CD patternsIdentifies common CI/CD configuration changes

Latest Papers

What's happening recently
View more

This work addresses customs clearance delays in global trade caused by ambiguous product descriptions and frequent updates to Harmonized System (HS) codes. To tackle this challenge, the authors propose a serverless MLOps framework that leverages event-driven pipelines and managed services to enable end-to-end, model-agnostic machine learning lifecycle management. The architecture supports automatic scaling, reproducible training, auditable deployment, and automated A/B testing, ensuring secure and seamless model transitions. By integrating custom text embeddings with models such as Text-CNN, the system achieves 98% accuracy on real-world HS code prediction tasks, meeting stringent service-level agreement (SLA) requirements. This approach significantly reduces long-term operational costs and establishes an efficient, cost-effective, and reproducible deployment paradigm for industrial-scale machine learning systems.

Harmonized System Code PredictionIndustrial Machine LearningMLOps

This study addresses the prevailing gap in AI education, which emphasizes model development while neglecting system engineering practices, leaving students ill-equipped to handle real-world challenges such as architectural design, deployment, and monitoring. To bridge this gap, the authors implemented a master’s-level course in which students built a movie recommendation system under realistic constraints, with a focus on integrating AI components into robust software systems, adopting data-driven machine learning practices, and cultivating systems-level thinking. Using a mixed-methods approach—combining analysis of student project artifacts with survey data—the research evaluates learners’ performance in architectural decision-making, integration of heterogeneous models, and adaptation to evolving requirements. Findings reveal common difficulties students encounter in AI system engineering and demonstrate the course’s effectiveness in addressing critical deficiencies in AI engineering education and enhancing systems-aware competencies.

AI-enabled systemsarchitectural designmachine learning integration

This study addresses the prevalent ad hoc and non-standardized practices in model integration and deployment within MLOps projects, which often stem from a lack of systematic architectural guidance. To bridge this gap, the authors conduct a gray literature review of 103 online sources and apply thematic analysis to derive, for the first time, 25 architecturally significant best practices. These practices are systematically categorized into five thematic groups, with explicit articulation of each practice’s impact on overall system architecture. The resulting framework offers a structured, actionable set of guidelines for MLOps model integration and deployment, providing both researchers and engineering teams with a coherent theoretical foundation and practical reference for designing robust, scalable machine learning systems.

architectural guidancegray literature reviewMLOps

This work addresses the challenges of SLO violations and resource inefficiency in machine learning model serving caused by inadequate capacity planning. To this end, the authors propose an adaptive, feedback-driven load testing framework that formalizes the ML serving load testing process for the first time. The framework incorporates real-traffic-based workload calibration and a warm-up mechanism, combined with adaptive search, performance signal feedback control, convergence detection, and GPU monitoring to efficiently estimate the maximum sustainable throughput under SLO constraints. Evaluation across 14 industrial cases demonstrates that the approach reduces capacity estimation error from approximately 30% to 2–6%, with the warm-up mechanism improving accuracy by 22.2%. This significantly mitigates deployment incidents and enhances GPU resource utilization efficiency.

capacity planningload testingML model serving

Hot Scholars

MI

Michael I. Jordan

Professor of Electrical Engineering and Computer Sciences and Professor of Statistics, UC Berkeley
machine learningcomputer sciencestatisticsartificial intelligence
AP

Annibale Panichella

Associate Professor, Delft University of Technology
Software TestingSE4AITest GenerationSBSE
YX

Yijia Xiao

University of California, Los Angeles
AI for FinanceAgentsAI for ScienceMultimodal LLM
SG

Stefan Gugler

Postdoc at TU Berlin
Machine Learning for Quantum ChemistryTheoretical Chemistry
KV

Koushik Viswanadha

Unknown affiliation
Natural Language ProcessingNatural Language Understanding