Score
Designs and implements systems and pipelines that record, version, and manage experiment metadata (configurations, code, datasets), artifacts, metrics and lineage for machine-learning experiments to enable reproducibility, comparison, monitoring, analytics and governance. Builds automation and scalable infrastructure (tracking servers, storage, UIs, logging integrations and orchestration) to support experiment iteration, automated pipelines, monitoring and experimental data handling, including integrations with tools such as MLflow.
Massive, dynamic data streams in digital platforms render conventional ML monitoring methods ineffective or prohibitively costly in manual effort, forcing enterprises to downgrade to simpler models. Method: This paper proposes the Machine Learning Monitoring Agent (MLMA) framework, introducing a test-driven, automated retraining mechanism based on data-adaptive reference loss batches—designed to enable efficient closed-loop operations while preserving human-in-the-loop collaborative governance. The approach integrates design science principles, dynamic reference loss computation, key metric visualization, and human–AI collaborative workflows. Contribution/Results: Evaluated on a large-scale instant-delivery platform, MLMA supports concurrent monitoring of hundreds of models, significantly reduces manual intervention frequency, and sustains long-term online model performance stability. Its core contribution lies in unifying dynamic data adaptation, automated trigger logic, and human–AI collaboration—thereby overcoming critical technical bottlenecks in real-time monitoring and adaptive maintenance of large-scale ML systems.
Medical AI deployment is hindered by insufficient production readiness of machine learning (ML) training pipelines. Method: This paper presents a progressive architectural evolution path—monolithic (chaotic) → modular monolithic → microservices—using SPIRA, a voice-based pre-diagnostic system for respiratory insufficiency, as a case study. It systematically introduces continuous training (CT) and a software-quality-attribute-driven MLOps governance framework tailored to healthcare, integrating modular design, microservice decomposition, and engineered CI/CD pipelines. Contribution/Results: The approach significantly improves pipeline maintainability, fault tolerance, and scalability, enabling stable, iterative evolution of SPIRA. It establishes an “agile ML + robust software engineering” co-design paradigm, delivering a reusable methodology and practical benchmark for engineering medical AI in highly regulated environments.
The widespread adoption of open-source machine learning (ML) datasets and models has intensified risks including data poisoning, supply-chain attacks, and regulatory non-compliance. Method: This paper proposes the first verifiable, end-to-end ML provenance framework integrating Trusted Execution Environments (TEEs) and transparent logging—built upon SPDX/SLSA standards and leveraging Intel SGX, hash-chain-based immutable logging, and zero-knowledge proofs. Contribution/Results: The framework enables provable artifact authenticity, auditable end-to-end lineage, and co-guaranteed confidentiality and integrity—without compromising intellectual property rights over data or models. Evaluated on two real-world ML pipelines, it achieves 100% metadata tampering detection, full verifiable traceability from training to deployment, and negligible runtime overhead—demonstrating practical viability for secure, compliant ML operations.
Current IDEs lack intelligent, end-to-end support for the machine learning (ML) lifecycle, while MLOps platforms remain decoupled from coding environments. To bridge this gap, we propose a novel large language model (LLM)-enhanced intelligent IDE paradigm that deeply integrates LLMs into the development environment. This enables synergistic, closed-loop automation across code-level intelligent programming—such as code generation, debugging, and completion—and full-stack MLOps pipeline orchestration—including data validation, feature store management, data drift detection, retraining triggers, and CI/CD deployment. The system unifies development, experimentation, validation, and monitoring phases, significantly improving engineering efficiency and reproducibility. Empirical evaluation on the UCI Adult and M5 datasets demonstrates a 61% reduction in pipeline configuration time, a 45% improvement in experimental reproducibility, and a 14% increase in data drift detection accuracy.
Machine learning teams face significant challenges in multi-workflow concurrent environments, including redundant intermediate data storage, inefficient cross-pipeline sharing, and high collaboration overhead. To address these issues, this paper proposes and implements a data virtualization service architecture tailored for ML workflows. The architecture adopts a service-oriented design, integrating distributed data management with dynamic metadata mapping to enable logical abstraction, on-demand loading, and unified access to heterogeneous intermediate data. Compared to conventional materialized storage approaches, it reduces storage overhead by an average of 62% (measured empirically) and substantially decreases inter-team collaboration latency. The system has been deployed in production, stably supporting six ML applications and over thirty concurrent workflows, demonstrating linear scalability. This work establishes a lightweight, elastic, and reusable data virtualization paradigm for large-scale ML infrastructure.
This work addresses reproducibility challenges in collaborative machine learning, which often stem not merely from missing artifacts but from team misalignments in interpreting prior work, inconsistent component evolution, and difficulties in reconstructing experimental intent. To tackle these issues, the authors propose a novel two-layer socio-technical architecture that integrates interactive support into reproducibility frameworks for the first time. The lower layer leverages a data-centric ML management system with full lifecycle provenance tracking, while the upper layer employs an AI-mediated semantic interface to enable structured collaboration, explanatory discourse, and consensus building. Deployed over 19 months in a clinical research setting, the system effectively identified and mitigated persistent interaction breakdowns, substantially enhancing shared team understanding and experimental reproducibility.
This work addresses the heavy reliance on expert knowledge in designing and debugging scientific workflows, a challenge exacerbated by existing large language model approaches that directly generate code without ensuring transparency, reproducibility, or seamless system integration. To overcome these limitations, we propose an AI-assisted scientific workflow management framework that decouples user intent from implementation through a structured specification phase, enabling specification-driven workflow generation and validation. We further introduce a multi-layer debugging agent powered by large language models to automate error diagnosis and correction. By deeply integrating with the Pegasus workflow system via the Model Context Protocol (MCP), our approach supports end-to-end workflow lifecycle management. Empirical evaluation demonstrates successful generation and execution of federated learning medical imaging workflows comprising thousands of tasks, substantially reducing debugging effort and empowering non-expert users to construct complex workflows adhering to expert-level design patterns.
This work addresses a critical limitation in current AI4Science practices, which often treat datasets as static interfaces while neglecting the uncertainties and implicit assumptions introduced by the multi-stage processing pipeline from raw measurements to curated datasets. To remedy this, the paper proposes a “computable observation framework” that explicitly models this pipeline as an auditable and reproducible inference component, capturing its configuration, validity, and associated uncertainties. By integrating scientific workflow analysis, uncertainty quantification, and cross-dataset stability assessment, the framework enables the construction of domain-specific observation protocols. Empirical evaluation on large-scale neuroscience data reveals that only approximately 0.0004% of processing pipelines exhibit cross-dataset stability, exposing severe fragility in current practices and underscoring the framework’s essential role in uncovering hidden assumptions, validating transferability, and controlling for multiplicity.
This study addresses the widespread neglect of licensing terms and regulatory compliance in the deployment of machine learning models within open-source software, particularly in safety-critical contexts where associated risks are pronounced. The authors present the first systematic investigation of ML usage across 173 open-source projects on GitHub spanning 16 application domains. Through code inspection and contextual analysis, they evaluate each model’s role in decision-making, the presence of risk-mitigation strategies, and adherence to licensing requirements. The findings reveal that certain projects employ ML for high-stakes decisions without complying with applicable license conditions and often lack essential post-processing safeguards. This work uncovers critical compliance blind spots in the open-source ecosystem and provides an empirical foundation for developing compliance guidelines and automated detection tools.
This study addresses the unclear usage patterns and functional demands surrounding current MLOps frameworks in open-source projects, which hinder their effective evolution. For the first time, it systematically links real-world framework adoption with user enhancement requests by analyzing GitHub dependencies, API invocations, and issue reports across eight prominent MLOps frameworks, employing qualitative coding and thematic mapping. The findings reveal that developers prefer customized integrations over out-of-the-box solutions, and that these frameworks are seldom directly embedded in GitHub Workflows, instead being primarily applied to core machine learning phases and infrastructure governance. Users most frequently request enhancements to core functionality, greater API exposure, and improved CI/CD integration, while increasingly adopting multiple frameworks in tandem.