Score
Designs and builds modular, multi‑stage model training and evaluation pipelines, defining pipeline architecture, stages, integrations, generation, deployment, and management processes (including orchestration frameworks such as Kubeflow Pipelines). Analyzes and models pipeline dependencies, data movement, and execution flows, and implements profiling, forecasting, evaluation, and monitoring tooling to optimize, maintain, and integrate pipeline components.
To address the lack of a unified knowledge framework in MLOps, this paper conducts a multi-source literature review (MLR), systematically synthesizing 150 academic publications and 48 grey literature sources to overcome single-perspective limitations. Through thematic coding and cross-source evidence triangulation, it establishes the first comprehensive MLOps conceptual model and practice map spanning the full ML lifecycle and integrating consensus from both industry and academia. Key contributions include: (1) a widely adopted, rigorous definition of MLOps; (2) distillation of 12 core MLOps practices; and (3) identification of seven recurrent implementation challenges alongside empirically grounded mitigation strategies. The resulting knowledge base is modular, reusable, and rigorously validated—serving as a foundational reference for MLOps standardization, tooling development, and empirical research.
This work addresses the growing complexity of CI/CD pipelines and the lack of structured analysis capabilities in existing tools for understanding their behavior, failures, and version evolution. The authors propose an innovative approach that uniquely integrates digital twin technology with BPMN-based modeling in DevOps contexts. By automatically parsing raw CI configurations and execution logs, the method constructs structured, high-level process models that enable pipeline visualization, failure traceability, and cross-version comparison. Evaluated across multiple open-source projects, the approach demonstrates effectiveness in monitoring, evolutionary analysis, and fault diagnosis, offering a modular and extensible foundational framework for the analysis and optimization of CI/CD pipelines.
Medical AI deployment is hindered by insufficient production readiness of machine learning (ML) training pipelines. Method: This paper presents a progressive architectural evolution path—monolithic (chaotic) → modular monolithic → microservices—using SPIRA, a voice-based pre-diagnostic system for respiratory insufficiency, as a case study. It systematically introduces continuous training (CT) and a software-quality-attribute-driven MLOps governance framework tailored to healthcare, integrating modular design, microservice decomposition, and engineered CI/CD pipelines. Contribution/Results: The approach significantly improves pipeline maintainability, fault tolerance, and scalability, enabling stable, iterative evolution of SPIRA. It establishes an “agile ML + robust software engineering” co-design paradigm, delivering a reusable methodology and practical benchmark for engineering medical AI in highly regulated environments.
This paper addresses the conceptual ambiguity, ill-defined boundaries, and lack of implementation standards between Infrastructure-as-Code (IaC) and Pipeline-as-Code in DevOps practice. To resolve these issues, we systematically delineate their respective roles and synergistic mechanisms within the DevOps ecosystem and propose a reusable, standardized IaC-driven CI/CD implementation framework. Our approach integrates Terraform for infrastructure provisioning, Ansible for configuration management, GitLab CI for pipeline orchestration, and Docker/Kubernetes for containerized deployment—enabling an end-to-end automated delivery pipeline. Empirical evaluation demonstrates 99.8% configuration change accuracy, reduces environment provisioning time from hours to minutes, and significantly improves deployment consistency and delivery efficiency.
Existing software architecture frameworks inadequately model machine learning (ML) systems, as they overlook the needs of emerging stakeholders—such as data scientists and data engineers—and lack expressive support for ML-specific characteristics, including component uncertainty, heterogeneity, and collaborative behavior. Method: Through an empirical study involving interviews and surveys with 61 domain experts from 25 organizations across 10 countries, we systematically identified ML-relevant stakeholders and their concerns for the first time. Contribution/Results: We propose novel, ML-adapted architectural viewpoints and views, extending traditional frameworks to enable unified modeling of both ML and non-ML components. This yields the *ML-Enhanced Systems Architecture Framework Extension Guide*, which has been preliminarily adopted in industry for intelligent system architecture governance. Our work bridges a critical theoretical and practical gap in stakeholder modeling and viewpoint systematization for ML system architecture design.
This work addresses the inadequacy of existing large language model (LLM) lifecycle frameworks, which predominantly emphasize operational efficiency while lacking explicit support for security-critical activities—such as data provenance, component signing, and access control—and failing to align governance requirements with specific lifecycle phases. The paper proposes the first security-oriented LLM system lifecycle model, structured not by workflow but by security boundaries, organizing 32 phases into four layered pipelines: data, model, distribution, and application, while integrating LLMOps and governance pillars. It uniquely identifies 13 distinct security-critical phases and exposes a structural imbalance wherein regulatory evidence is concentrated at deployment despite pivotal decisions occurring during development. By mapping key standards—including NIST AI RMF, the EU AI Act, and ISO/IEC 42001—the study establishes a phase-to-governance correspondence mechanism, yielding a comprehensive, lifecycle-spanning security analysis framework that offers structured guidance for compliance and secure design.
This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.
Community-driven scientific workflow ecosystems often struggle to sustain themselves due to ambiguous maintenance and user support mechanisms, particularly in cross-platform collaboration and heterogeneous execution environments. This study presents the first cross-platform empirical analysis of the nf-core ecosystem, systematically examining 15,760 GitHub issues, 35,411 pull requests, and 895 forum discussions. By integrating metadata and textual features into predictive models, the research uncovers significant disparities in maintenance and support activities across platforms and highlights weak explicit linkages among them. The findings reveal that issues, pull requests, and forum posts predominantly serve distinct roles—coordinating maintenance, facilitating code integration, and providing user support, respectively. Moreover, issue actionability, diagnostic evidence, and depth of interaction emerge as critical determinants of resolution efficiency.
This work addresses the heavy reliance on expert knowledge in designing and debugging scientific workflows, a challenge exacerbated by existing large language model approaches that directly generate code without ensuring transparency, reproducibility, or seamless system integration. To overcome these limitations, we propose an AI-assisted scientific workflow management framework that decouples user intent from implementation through a structured specification phase, enabling specification-driven workflow generation and validation. We further introduce a multi-layer debugging agent powered by large language models to automate error diagnosis and correction. By deeply integrating with the Pegasus workflow system via the Model Context Protocol (MCP), our approach supports end-to-end workflow lifecycle management. Empirical evaluation demonstrates successful generation and execution of federated learning medical imaging workflows comprising thousands of tasks, substantially reducing debugging effort and empowering non-expert users to construct complex workflows adhering to expert-level design patterns.
This work proposes RefineGPT, a domain-specific intelligent agent for the automated synthesis of unit-level process diagrams (UPDs) in petroleum refining, addressing the semantic gap between natural language design intents and the rigorous physical logic of refining engineering. The approach employs a hierarchical architecture wherein a small language model selects appropriate process units based on design requirements, while a large language model constructs the complete process topology. Innovatively, the method introduces a pipeline that extracts implicit process patterns from unstructured historical diagrams to synthesize high-quality, chain-of-reasoning–based training data, thereby deeply integrating domain knowledge with large-model capabilities. Experimental results demonstrate that RefineGPT significantly outperforms existing methods in both topological consistency and chemical feasibility, offering a high-fidelity pathway toward AI-driven industrial process synthesis.