ml infrastructure

Designs and builds scalable software systems and platforms that enable the development, training, deployment, and operation of machine learning and AI models, including data pipelines, feature stores, training and distributed compute platforms, model/experiment registries, serving/inference layers, and monitoring/observability. Analyzes and optimizes the performance, reliability, cost, security, and automation (e.g., CI/CD, orchestration) of end-to-end ML workflows and infrastructure components.

mlinfrastructure

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$228K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Is Your Training Pipeline Production-Ready? A Case Study in the Healthcare Domain

Jun 07, 2025
DL
Daniel Lawand
🏛️ University of São Paulo | Tilburg University | Technical University of Eindhoven

Medical AI deployment is hindered by insufficient production readiness of machine learning (ML) training pipelines. Method: This paper presents a progressive architectural evolution path—monolithic (chaotic) → modular monolithic → microservices—using SPIRA, a voice-based pre-diagnostic system for respiratory insufficiency, as a case study. It systematically introduces continuous training (CT) and a software-quality-attribute-driven MLOps governance framework tailored to healthcare, integrating modular design, microservice decomposition, and engineered CI/CD pipelines. Contribution/Results: The approach significantly improves pipeline maintainability, fault tolerance, and scalability, enabling stable, iterative evolution of SPIRA. It establishes an “agile ML + robust software engineering” co-design paradigm, delivering a reusable methodology and practical benchmark for engineering medical AI in highly regulated environments.

Ensuring ML training pipelines are production-ready in healthcareEvolving architecture for better maintainability and robustnessImproving software quality in MLES for respiratory pre-diagnosis

A Large-Scale Study of Model Integration in ML-Enabled Software Systems

Aug 12, 2024
YS
Yorick Sens
🏛️ Ruhr University Bochum | Chalmers University of Gothenburg

This study addresses the challenges of integrating machine learning (ML) models into software systems—namely, poor integration practices, low reusability, and unclear architectural boundaries. It presents the first large-scale empirical investigation across 2,928 open-source ML-enabled systems. Leveraging GitHub code mining, static analysis, topic modeling, and architectural pattern identification, the work systematically characterizes ML integration topologies, code/model reuse practices, and maintenance bottlenecks. Key contributions include: (1) the first comprehensive classification framework and architectural pattern atlas for ML-enabled systems; (2) identification of seven prevalent integration topologies and four model reuse patterns; and (3) uncovering critical interdisciplinary collaboration barriers in ML-software co-development. The findings bridge the methodological gap between data science and software engineering at the model embedding stage, providing industry-practical architectural guidelines that significantly enhance the maintainability and reusability of ML systems.

ML and code reuse practicesML model integration challengesML-enabled system characteristics

MLOps Monitoring at Scale for Digital Platforms

Apr 23, 2025
YJ
Yu Jeffrey Hu
🏛️ Purdue University | Essec Business School | Maastricht University

Massive, dynamic data streams in digital platforms render conventional ML monitoring methods ineffective or prohibitively costly in manual effort, forcing enterprises to downgrade to simpler models. Method: This paper proposes the Machine Learning Monitoring Agent (MLMA) framework, introducing a test-driven, automated retraining mechanism based on data-adaptive reference loss batches—designed to enable efficient closed-loop operations while preserving human-in-the-loop collaborative governance. The approach integrates design science principles, dynamic reference loss computation, key metric visualization, and human–AI collaborative workflows. Contribution/Results: Evaluated on a large-scale instant-delivery platform, MLMA supports concurrent monitoring of hundreds of models, significantly reduces manual intervention frequency, and sustains long-term online model performance stability. Its core contribution lies in unifying dynamic data adaptation, automated trigger logic, and human–AI collaboration—thereby overcoming critical technical bottlenecks in real-time monitoring and adaptive maintenance of large-scale ML systems.

Automating re-training to maintain model performance at scaleMonitoring ML models in large unstable data streamsReducing labor-intensive MLOps supervision in digital platforms

This work addresses customs clearance delays in global trade caused by ambiguous product descriptions and frequent updates to Harmonized System (HS) codes. To tackle this challenge, the authors propose a serverless MLOps framework that leverages event-driven pipelines and managed services to enable end-to-end, model-agnostic machine learning lifecycle management. The architecture supports automatic scaling, reproducible training, auditable deployment, and automated A/B testing, ensuring secure and seamless model transitions. By integrating custom text embeddings with models such as Text-CNN, the system achieves 98% accuracy on real-world HS code prediction tasks, meeting stringent service-level agreement (SLA) requirements. This approach significantly reduces long-term operational costs and establishes an efficient, cost-effective, and reproducible deployment paradigm for industrial-scale machine learning systems.

Harmonized System Code PredictionIndustrial Machine LearningMLOps

Enhancing Architecture Frameworks by Including Modern Stakeholders and their Views/Viewpoints

Aug 09, 2023
AM
Armin Moin
🏛️ University of Colorado | University of Reading | Technical University of Munich | University of Antwerp

Existing software architecture frameworks inadequately model machine learning (ML) systems, as they overlook the needs of emerging stakeholders—such as data scientists and data engineers—and lack expressive support for ML-specific characteristics, including component uncertainty, heterogeneity, and collaborative behavior. Method: Through an empirical study involving interviews and surveys with 61 domain experts from 25 organizations across 10 countries, we systematically identified ML-relevant stakeholders and their concerns for the first time. Contribution/Results: We propose novel, ML-adapted architectural viewpoints and views, extending traditional frameworks to enable unified modeling of both ML and non-ML components. This yields the *ML-Enhanced Systems Architecture Framework Extension Guide*, which has been preliminarily adopted in industry for intelligent system architecture governance. Our work bridges a critical theoretical and practical gap in stakeholder modeling and viewpoint systematization for ML system architecture design.

Architectural FrameworksDesign MethodologyMachine Learning Systems

Latest Papers

What's happening recently
View more

This work addresses the inefficiencies in large-scale recommendation systems caused by maintaining separate models for different scenarios and objectives, which hinders development velocity and delays technology adoption. To overcome this, the authors propose the Standardized Model Template (SMT) framework, which leverages composable, standardized machine learning components to enable “design once, deploy everywhere,” uniformly accommodating diverse data distributions and optimization objectives. By decoupling model architecture from scenario-specific configurations, SMT reduces the complexity of technology deployment from O(n·2ᵏ) to O(n+k), breaking away from the conventional “one objective, one model” paradigm. Empirical evaluation on Meta’s ad ranking system demonstrates that SMT improves average cross-entropy by 0.63%, reduces engineering time per model iteration by 92%, and increases the throughput of technology-model pair adoption by 6.3×.

computational advertisinglarge-scale ML ecosystemsML technique propagation

This study addresses the prevalent ad hoc and non-standardized practices in model integration and deployment within MLOps projects, which often stem from a lack of systematic architectural guidance. To bridge this gap, the authors conduct a gray literature review of 103 online sources and apply thematic analysis to derive, for the first time, 25 architecturally significant best practices. These practices are systematically categorized into five thematic groups, with explicit articulation of each practice’s impact on overall system architecture. The resulting framework offers a structured, actionable set of guidelines for MLOps model integration and deployment, providing both researchers and engineering teams with a coherent theoretical foundation and practical reference for designing robust, scalable machine learning systems.

architectural guidancegray literature reviewMLOps

This study addresses the prevailing gap in AI education, which emphasizes model development while neglecting system engineering practices, leaving students ill-equipped to handle real-world challenges such as architectural design, deployment, and monitoring. To bridge this gap, the authors implemented a master’s-level course in which students built a movie recommendation system under realistic constraints, with a focus on integrating AI components into robust software systems, adopting data-driven machine learning practices, and cultivating systems-level thinking. Using a mixed-methods approach—combining analysis of student project artifacts with survey data—the research evaluates learners’ performance in architectural decision-making, integration of heterogeneous models, and adaptation to evolving requirements. Findings reveal common difficulties students encounter in AI system engineering and demonstrate the course’s effectiveness in addressing critical deficiencies in AI engineering education and enhancing systems-aware competencies.

AI-enabled systemsarchitectural designmachine learning integration

This work addresses the inefficiency of the current Python machine learning ecosystem in supporting large-scale, highly concurrent machine learning pipeline searches driven by large language model (LLM) agents. To overcome this limitation, we propose the first system architecture specifically designed for agent-driven ML workloads, which decouples the agent’s planning and reasoning from pipeline execution to enable batch compilation and efficient scheduling. Our system introduces pipeline graph compilation, batched execution optimization, and a high-performance Rust runtime, while seamlessly integrating with mainstream Python libraries and supporting heterogeneous backends including CPUs and GPUs. Experimental results demonstrate that our approach achieves up to a 16.6× speedup on large-scale agent-driven ML pipeline search tasks.

agentic pipeline searchlarge language modelsmassive ML workloads

This work addresses the challenges of configuring and tuning the complex Linux kernel, where deploying machine learning (ML) models directly in kernel space is hindered by the absence of floating-point unit (FPU) support and prohibitive performance overhead. To overcome these limitations, the authors propose the first lightweight ML infrastructure tailored for the Linux kernel, enabling safe and efficient model inference without FPU usage through a cooperative kernel–user space design. The architecture comprises a kernel module, a lightweight inference agent, and a cross-space communication interface. A prototype implementation demonstrates the feasibility of this approach, and experimental results show that it introduces ML capabilities into the kernel with low overhead and high scalability, opening a new avenue for intelligent kernel optimization.

floating-point operationskernel spaceLinux kernel

Hot Scholars

JZ

Jinhua Zhu

University of Science and Technology of China
Machine Learning
WZ

Wengang Zhou

Professor, EEIS Department, University of Science and Technology of China
Multimedia RetrievalComputer VisionComputer Game
HL

Houqiang Li

Professor, Department of Electric Engineering and Information Science, University of Science and
Multimedia SearchImage/Video AnalysisImage/Video Coding