Score
Designs, implements, and evaluates software frameworks and libraries that support the end‑to‑end development, training, evaluation, and deployment of machine learning models. This includes API and abstraction design, modular training/evaluation pipelines, model serialization and serving, integration with hardware accelerators and distributed systems, tooling for data I/O, hyperparameter tuning, and runtime monitoring and performance optimization.
To address low testing and debugging efficiency and immature toolchains when deploying rapidly evolving large language models (LLMs) on emerging platforms (e.g., browsers, mobile devices), this paper proposes TapML—a top-down, test-driven framework. Methodologically, TapML introduces (1) the first operator-level test pruning technique that automatically generates high-coverage, realistic test inputs; (2) a progressive cross-platform migration strategy that significantly narrows the scope for compound error localization; and (3) native backend support for Metal and WebGPU, with deep integration into MLC-LLM. Evaluated over two years, TapML has enabled efficient deployment of 105 emerging models—spanning 27 distinct architectures—across five platform categories, reducing average deployment time by 42%. It has since become the default development paradigm for MLC-LLM.
Small- and medium-sized enterprises (SMEs) face prohibitive costs, high production downtime risks, and operational complexity when integrating machine learning (ML) into legacy industrial systems. Method: This paper proposes a human-in-the-loop interactive ML framework that decouples the ML model lifecycle from the production environment via an API-based middleware layer. It employs a lightweight model-serving architecture and a browser-based interactive interface, enabling zero-hardware-upgrade deployment, zero-downtime integration, and remote real-time parameter tuning. Contribution/Results: The framework is the first to support dynamic model maintenance and online collaborative decision-making by domain experts—without modifying existing systems. Experimental evaluation demonstrates substantial reductions in ML adoption barriers and implementation costs, alongside measurable improvements in manufacturing quality and safety. The solution exhibits strong scalability and engineering practicality for industrial deployment.
This study addresses the challenges of integrating machine learning (ML) models into software systems—namely, poor integration practices, low reusability, and unclear architectural boundaries. It presents the first large-scale empirical investigation across 2,928 open-source ML-enabled systems. Leveraging GitHub code mining, static analysis, topic modeling, and architectural pattern identification, the work systematically characterizes ML integration topologies, code/model reuse practices, and maintenance bottlenecks. Key contributions include: (1) the first comprehensive classification framework and architectural pattern atlas for ML-enabled systems; (2) identification of seven prevalent integration topologies and four model reuse patterns; and (3) uncovering critical interdisciplinary collaboration barriers in ML-software co-development. The findings bridge the methodological gap between data science and software engineering at the model embedding stage, providing industry-practical architectural guidelines that significantly enhance the maintainability and reusability of ML systems.
Existing software architecture frameworks inadequately model machine learning (ML) systems, as they overlook the needs of emerging stakeholders—such as data scientists and data engineers—and lack expressive support for ML-specific characteristics, including component uncertainty, heterogeneity, and collaborative behavior. Method: Through an empirical study involving interviews and surveys with 61 domain experts from 25 organizations across 10 countries, we systematically identified ML-relevant stakeholders and their concerns for the first time. Contribution/Results: We propose novel, ML-adapted architectural viewpoints and views, extending traditional frameworks to enable unified modeling of both ML and non-ML components. This yields the *ML-Enhanced Systems Architecture Framework Extension Guide*, which has been preliminarily adopted in industry for intelligent system architecture governance. Our work bridges a critical theoretical and practical gap in stakeholder modeling and viewpoint systematization for ML system architecture design.
Medical AI deployment is hindered by insufficient production readiness of machine learning (ML) training pipelines. Method: This paper presents a progressive architectural evolution path—monolithic (chaotic) → modular monolithic → microservices—using SPIRA, a voice-based pre-diagnostic system for respiratory insufficiency, as a case study. It systematically introduces continuous training (CT) and a software-quality-attribute-driven MLOps governance framework tailored to healthcare, integrating modular design, microservice decomposition, and engineered CI/CD pipelines. Contribution/Results: The approach significantly improves pipeline maintainability, fault tolerance, and scalability, enabling stable, iterative evolution of SPIRA. It establishes an “agile ML + robust software engineering” co-design paradigm, delivering a reusable methodology and practical benchmark for engineering medical AI in highly regulated environments.
This study addresses the prevalent ad hoc and non-standardized practices in model integration and deployment within MLOps projects, which often stem from a lack of systematic architectural guidance. To bridge this gap, the authors conduct a gray literature review of 103 online sources and apply thematic analysis to derive, for the first time, 25 architecturally significant best practices. These practices are systematically categorized into five thematic groups, with explicit articulation of each practice’s impact on overall system architecture. The resulting framework offers a structured, actionable set of guidelines for MLOps model integration and deployment, providing both researchers and engineering teams with a coherent theoretical foundation and practical reference for designing robust, scalable machine learning systems.
This study addresses the widespread neglect of licensing terms and regulatory compliance in the deployment of machine learning models within open-source software, particularly in safety-critical contexts where associated risks are pronounced. The authors present the first systematic investigation of ML usage across 173 open-source projects on GitHub spanning 16 application domains. Through code inspection and contextual analysis, they evaluate each model’s role in decision-making, the presence of risk-mitigation strategies, and adherence to licensing requirements. The findings reveal that certain projects employ ML for high-stakes decisions without complying with applicable license conditions and often lack essential post-processing safeguards. This work uncovers critical compliance blind spots in the open-source ecosystem and provides an empirical foundation for developing compliance guidelines and automated detection tools.
This study addresses the lack of systematic empirical investigation into the deployment of large language model (LLM) inference frameworks in real-world software systems. Through large-scale analysis of open-source projects, combining static code analysis with metadata mining, this work systematically characterizes the adoption patterns and integration strategies of prominent frameworks—including vLLM, SGLang, TensorRT-LLM, LMDeploy, and FlashInfer—and examines their relationships with model type, scale, modality, and deployment environment. The findings reveal that vLLM exhibits the highest adoption rate, while multi-framework co-deployment remains limited yet holds complementary potential. Framework selection is strongly driven by model characteristics and application scenarios, effectively supporting diverse system designs such as reinforcement learning inference, multimodal generation, and microservice architectures.
This study addresses the prevailing gap in AI education, which emphasizes model development while neglecting system engineering practices, leaving students ill-equipped to handle real-world challenges such as architectural design, deployment, and monitoring. To bridge this gap, the authors implemented a master’s-level course in which students built a movie recommendation system under realistic constraints, with a focus on integrating AI components into robust software systems, adopting data-driven machine learning practices, and cultivating systems-level thinking. Using a mixed-methods approach—combining analysis of student project artifacts with survey data—the research evaluates learners’ performance in architectural decision-making, integration of heterogeneous models, and adaptation to evolving requirements. Findings reveal common difficulties students encounter in AI system engineering and demonstrate the course’s effectiveness in addressing critical deficiencies in AI engineering education and enhancing systems-aware competencies.
This work proposes the first end-to-end automated artificial intelligence research framework capable of fully automating the development pipeline from algorithmic idea generation to executable machine learning classifiers. The approach integrates structured meta-prompt engineering with large language model–based code generation, augmented by an automated evaluation and iterative optimization mechanism. Experimental results on twenty standard datasets from the Infinity-Bench benchmark demonstrate that multiple novel classifiers autonomously generated by the framework significantly outperform baseline methods implemented in scikit-learn. This study thus achieves, for the first time, complete automation of the entire workflow—from initial algorithmic conception to deployable, runnable code—marking a significant step toward self-driving AI research systems.