Score
Designs and implements end-to-end processes and systems that take models from data and training through validation, deployment, inference serving, and eventual retirement, including pipelines, automation, and serving infrastructure. Defines and operates versioning, reproducibility, monitoring, alerting, drift detection, retraining triggers, and governance to ensure model performance, correctness, scalability, and maintainability in production.
To address the lack of a unified knowledge framework in MLOps, this paper conducts a multi-source literature review (MLR), systematically synthesizing 150 academic publications and 48 grey literature sources to overcome single-perspective limitations. Through thematic coding and cross-source evidence triangulation, it establishes the first comprehensive MLOps conceptual model and practice map spanning the full ML lifecycle and integrating consensus from both industry and academia. Key contributions include: (1) a widely adopted, rigorous definition of MLOps; (2) distillation of 12 core MLOps practices; and (3) identification of seven recurrent implementation challenges alongside empirically grounded mitigation strategies. The resulting knowledge base is modular, reusable, and rigorously validated—serving as a foundational reference for MLOps standardization, tooling development, and empirical research.
To address challenges in Cyber-Physical Systems (CPS) development—including heterogeneous formal models, fragmented storage of modeling artifacts, inadequate version management, and limited knowledge reuse—this paper proposes an ontology-driven engineering knowledge graph framework. It introduces a unified systems engineering ontology built upon the custom Ontology Modelling Language (OML), enabling semantic integration of modeling artifacts across formal methods (e.g., SysML, UML, Modelica). The framework integrates a workflow engine, SPARQL querying, SWRL rule-based reasoning, and versioned graph storage to implicitly encapsulate complex knowledge graph operations. It is the first to support full-lifecycle semantic interoperability and automated knowledge discovery. Evaluated on an electric-drive intelligent sensor system, the framework significantly improves model version management efficiency, accelerates information retrieval, and uncovers three categories of latent engineering knowledge via inference.
Medical AI deployment is hindered by insufficient production readiness of machine learning (ML) training pipelines. Method: This paper presents a progressive architectural evolution path—monolithic (chaotic) → modular monolithic → microservices—using SPIRA, a voice-based pre-diagnostic system for respiratory insufficiency, as a case study. It systematically introduces continuous training (CT) and a software-quality-attribute-driven MLOps governance framework tailored to healthcare, integrating modular design, microservice decomposition, and engineered CI/CD pipelines. Contribution/Results: The approach significantly improves pipeline maintainability, fault tolerance, and scalability, enabling stable, iterative evolution of SPIRA. It establishes an “agile ML + robust software engineering” co-design paradigm, delivering a reusable methodology and practical benchmark for engineering medical AI in highly regulated environments.
This work addresses the challenge of reliably translating natural language into industrial-grade, deployable SysMLv2 models. The authors propose an iterative generate-check-repair framework that, for the first time, integrates a production-level SysMLv2 conformance checker directly into the generation process as a control mechanism rather than a post-processing step. By combining large language model (LLM) generation with deterministic diagnostic feedback and targeted repair strategies—and terminating only when zero errors remain—the method achieves perfect compliance. Evaluated across 604 test cases derived from 151 prompts and four distinct LLMs, the approach elevates single-pass generation compliance from 51.16% to 100%, enabling robust, direct translation of natural language specifications into engineering-ready SysMLv2 models.
Existing conformance checking approaches between process models and reference models suffer from limited semantic expressiveness and insufficient automation, hindering fine-grained compliance verification. This paper proposes a semantic consistency checking method grounded in causal dependency analysis of tasks and events, transcending traditional trajectory-based dependency modeling by formally encoding causal constraints at the semantic level. We establish a unified framework integrating causal dependency modeling, semantic representation, and formal verification, and design an automated conformance checking algorithm implemented in a prototype tool. Empirical evaluation demonstrates that our approach significantly outperforms state-of-the-art techniques in both accuracy and flexibility, achieving— for the first time—the fully automated, high-expressivity semantic conformance verification of process models against reference models.
This work addresses the lack of end-to-end traceability from high-level models to generated code in Model-Driven Engineering (MDE) by proposing the ProMoTA framework. ProMoTA unifies the entire modeling and code generation process—spanning platform-independent models, platform-specific models, and final code—through megamodels and model transformation chains. The framework innovatively extends the Acceleo language to support fine-grained local traceability and, for the first time, enables comprehensive global traceability mapping and analysis across the full MDE lifecycle. Implemented on the Eclipse platform, ProMoTA’s effectiveness in facilitating end-to-end traceability analysis is empirically validated through a case study in wireless sensor network-based Internet of Things applications.
This study addresses the challenge faced by production system engineers in automatically verifying production line layouts due to limited knowledge of PDDL and planning theory. To bridge this gap, the authors propose a novel approach based on an Asset Administration Shell (AAS) capability model that natively generates complete PDDL planning problems directly from domain-level descriptions, eliminating the need for PDDL-specific submodels. The method integrates four Industry 4.0 standards—VDI 3682, IEC 61360-1, IDTA 02011, and IDTA 02016—to construct the AAS and employs an extraction algorithm to automatically translate multi-AAS architectures into PDDL domains. In a laboratory case study, the approach enabled engineers to systematically compare four layout variants by modifying only the AAS model, significantly lowering the barrier to adopting automated planning in industrial settings.
This work addresses the inadequacy of existing large language model (LLM) lifecycle frameworks, which predominantly emphasize operational efficiency while lacking explicit support for security-critical activities—such as data provenance, component signing, and access control—and failing to align governance requirements with specific lifecycle phases. The paper proposes the first security-oriented LLM system lifecycle model, structured not by workflow but by security boundaries, organizing 32 phases into four layered pipelines: data, model, distribution, and application, while integrating LLMOps and governance pillars. It uniquely identifies 13 distinct security-critical phases and exposes a structural imbalance wherein regulatory evidence is concentrated at deployment despite pivotal decisions occurring during development. By mapping key standards—including NIST AI RMF, the EU AI Act, and ISO/IEC 42001—the study establishes a phase-to-governance correspondence mechanism, yielding a comprehensive, lifecycle-spanning security analysis framework that offers structured guidance for compliance and secure design.
This work addresses the challenge of silent updates to large language models (LLMs) by service providers, which often occur without version changes and can lead to behavioral drift and functional regressions, while existing mechanisms lack deployment-side control over compatibility governance. Framing LLM updates as a software supply chain governance problem, this study proposes a deployment-side control framework that defines rule-based production contracts, constructs risk-category-oriented test suites, and enforces compatibility gates to validate model safety and performance prior to updates. Experimental results demonstrate that the approach effectively uncovers fine-grained regressions missed by aggregate metrics, while also highlighting critical challenges in test design, threshold calibration, and drift attribution.