ai platforms

Designs and implements systems that provide the infrastructure, orchestration, and developer-facing tools needed to train, deploy, monitor, and manage machine learning and AI models across their lifecycle. Work includes building scalable compute and data pipelines, model serving and APIs, experiment tracking, CI/CD for models, access controls, observability, and governance features.

aiplatforms

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.92
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$208K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Is Your Training Pipeline Production-Ready? A Case Study in the Healthcare Domain

Jun 07, 2025
DL
Daniel Lawand
🏛️ University of São Paulo | Tilburg University | Technical University of Eindhoven

Medical AI deployment is hindered by insufficient production readiness of machine learning (ML) training pipelines. Method: This paper presents a progressive architectural evolution path—monolithic (chaotic) → modular monolithic → microservices—using SPIRA, a voice-based pre-diagnostic system for respiratory insufficiency, as a case study. It systematically introduces continuous training (CT) and a software-quality-attribute-driven MLOps governance framework tailored to healthcare, integrating modular design, microservice decomposition, and engineered CI/CD pipelines. Contribution/Results: The approach significantly improves pipeline maintainability, fault tolerance, and scalability, enabling stable, iterative evolution of SPIRA. It establishes an “agile ML + robust software engineering” co-design paradigm, delivering a reusable methodology and practical benchmark for engineering medical AI in highly regulated environments.

Ensuring ML training pipelines are production-ready in healthcareEvolving architecture for better maintainability and robustnessImproving software quality in MLES for respiratory pre-diagnosis

This study addresses the prevailing gap in AI education, which emphasizes model development while neglecting system engineering practices, leaving students ill-equipped to handle real-world challenges such as architectural design, deployment, and monitoring. To bridge this gap, the authors implemented a master’s-level course in which students built a movie recommendation system under realistic constraints, with a focus on integrating AI components into robust software systems, adopting data-driven machine learning practices, and cultivating systems-level thinking. Using a mixed-methods approach—combining analysis of student project artifacts with survey data—the research evaluates learners’ performance in architectural decision-making, integration of heterogeneous models, and adaptation to evolving requirements. Findings reveal common difficulties students encounter in AI system engineering and demonstrate the course’s effectiveness in addressing critical deficiencies in AI engineering education and enhancing systems-aware competencies.

AI-enabled systemsarchitectural designmachine learning integration

Empirical Analysis on CI/CD Pipeline Evolution in Machine Learning Projects

Mar 18, 2024
AH
Alaa Houerbi
🏛️ University of Michigan- Dearborn

This study presents the first empirical investigation into the evolution of CI/CD configurations in machine learning (ML) projects. Addressing the lack of understanding regarding how CI/CD configurations co-evolve with ML components, the authors analyze 508 open-source ML projects, 343 manually annotated commits, and 15,634 automated CI/CD commits. They propose a novel 14-category taxonomy capturing synergistic changes between CI/CD and ML components, develop a dedicated clustering tool to identify recurrent evolutionary patterns, and establish an empirically grounded model linking developer experience to CI/CD configuration modification behavior. Results show that 61.8% of CI/CD-related commits involve build strategy modifications; common anti-patterns—including dependency hardcoding and missing test frameworks—are identified; and senior developers modify CI/CD configurations more frequently and effectively than juniors, confirming the critical role of experience in CI/CD maintenance.

Analyzes CI/CD evolution in ML projectsDevelops clustering tool for CI/CD patternsIdentifies common CI/CD configuration changes

To address the escalating computational demands, high costs, and cross-environment coordination challenges in large language model (LLM) training, this project develops an end-to-end hybrid-cloud AI infrastructure comprising the cloud-based multi-tenant supercomputing platform Vela and the on-premises ultra-large-scale training system Blue Vela. It introduces a novel cloud-edge collaborative dynamic resource scheduling paradigm, integrating AI-optimized hardware clusters, a full-stack software–hardware co-designed training framework, and a unified telemetry and elastic orchestration system. Compared to conventional approaches, the infrastructure achieves a 35% improvement in training efficiency at the thousand-GPU scale and reduces fault recovery time by 60%. It has successfully accelerated iterative development of IBM’s third-generation and beyond generative AI models, while delivering commercial inference services with millisecond-level latency and 99.99% system availability.

AI model developmentcomputational resourcestraining cost

Enhancing Architecture Frameworks by Including Modern Stakeholders and their Views/Viewpoints

Aug 09, 2023
AM
Armin Moin
🏛️ University of Colorado | University of Reading | Technical University of Munich | University of Antwerp

Existing software architecture frameworks inadequately model machine learning (ML) systems, as they overlook the needs of emerging stakeholders—such as data scientists and data engineers—and lack expressive support for ML-specific characteristics, including component uncertainty, heterogeneity, and collaborative behavior. Method: Through an empirical study involving interviews and surveys with 61 domain experts from 25 organizations across 10 countries, we systematically identified ML-relevant stakeholders and their concerns for the first time. Contribution/Results: We propose novel, ML-adapted architectural viewpoints and views, extending traditional frameworks to enable unified modeling of both ML and non-ML components. This yields the *ML-Enhanced Systems Architecture Framework Extension Guide*, which has been preliminarily adopted in industry for intelligent system architecture governance. Our work bridges a critical theoretical and practical gap in stakeholder modeling and viewpoint systematization for ML system architecture design.

Architectural FrameworksDesign MethodologyMachine Learning Systems

Latest Papers

What's happening recently
View more

This study addresses the disruptive impact of large language models and AI agent systems—capable of generating vast volumes of code—on traditional software engineering paradigms. The work proposes a new paradigm centered on agent orchestration, verification of AI-generated code, and structured human-AI collaboration. Through a structured synthesis of literature review and industry practices, it constructs a comprehensive framework encompassing education, toolchains, lifecycle management, and governance. The research reveals a fundamental shift in the nature of code—from a scarce craft artifact to a consumable commodity—and identifies the evolving role of software engineers toward system design, semantic validation, and accountability oversight. It further establishes key directions such as a verification-first software development lifecycle, offering both theoretical grounding and practical pathways for software engineering transformation in the AI era.

Agentic AI SystemsAI-generated CodeHuman-AI Collaboration

This study addresses the widespread absence of systematic instruction on building, testing, deploying, and maintaining AI/ML systems in current undergraduate software engineering (SE) curricula. It presents the first comprehensive delineation of core AI/ML topics essential for SE practice, integrating curriculum mapping analysis, instructor needs surveys, and structured modeling to identify critical content gaps in existing programs. Grounded in empirical evidence, the work proposes actionable pathways for embedding high-priority AI/ML themes into established SE courses. The resulting framework offers a practical, implementable guide to enhance SE education’s capacity to support the development of intelligent software systems.

AI/ML-based SoftwareCurriculum IntegrationMachine Learning