Score
Designs and builds software systems that incorporate AI/ML models to perform automated inference and decision-making, encompassing model training pipelines, data engineering, and deployment/inference infrastructure. Analyzes and maintains these systems' performance, scalability, reliability, safety, and lifecycle through evaluation metrics, monitoring, testing, versioning, and integration with other services.
Real-world deployment of AI systems faces dual challenges: heterogeneous data floods and stringent low-latency requirements—straining conventional software architectures to their limits. To address this, we conduct the first systematic analysis of 217 production-deployed ML systems from a Data-Oriented Architecture (DOA) perspective, uncovering implicit DOA design patterns—particularly in loose coupling, decentralization, and data-driven orchestration. Leveraging systematic literature review, architectural pattern extraction, and requirement-to-design mapping analysis, we synthesize the first empirically grounded DOA practice guide and open-challenge taxonomy tailored to real-world ML deployments. Furthermore, we propose actionable, reusable recommendations for ML system deployment and introduce a novel architecture evaluation framework. Collectively, these contributions bridge the critical knowledge gap between DOA theory and industrial engineering practice.
This study addresses the prevailing gap in AI education, which emphasizes model development while neglecting system engineering practices, leaving students ill-equipped to handle real-world challenges such as architectural design, deployment, and monitoring. To bridge this gap, the authors implemented a master’s-level course in which students built a movie recommendation system under realistic constraints, with a focus on integrating AI components into robust software systems, adopting data-driven machine learning practices, and cultivating systems-level thinking. Using a mixed-methods approach—combining analysis of student project artifacts with survey data—the research evaluates learners’ performance in architectural decision-making, integration of heterogeneous models, and adaptation to evolving requirements. Findings reveal common difficulties students encounter in AI system engineering and demonstrate the course’s effectiveness in addressing critical deficiencies in AI engineering education and enhancing systems-aware competencies.
Current IDEs lack intelligent, end-to-end support for the machine learning (ML) lifecycle, while MLOps platforms remain decoupled from coding environments. To bridge this gap, we propose a novel large language model (LLM)-enhanced intelligent IDE paradigm that deeply integrates LLMs into the development environment. This enables synergistic, closed-loop automation across code-level intelligent programming—such as code generation, debugging, and completion—and full-stack MLOps pipeline orchestration—including data validation, feature store management, data drift detection, retraining triggers, and CI/CD deployment. The system unifies development, experimentation, validation, and monitoring phases, significantly improving engineering efficiency and reproducibility. Empirical evaluation on the UCI Adult and M5 datasets demonstrates a 61% reduction in pipeline configuration time, a 45% improvement in experimental reproducibility, and a 14% increase in data drift detection accuracy.
This study addresses the challenges of integrating machine learning (ML) models into software systems—namely, poor integration practices, low reusability, and unclear architectural boundaries. It presents the first large-scale empirical investigation across 2,928 open-source ML-enabled systems. Leveraging GitHub code mining, static analysis, topic modeling, and architectural pattern identification, the work systematically characterizes ML integration topologies, code/model reuse practices, and maintenance bottlenecks. Key contributions include: (1) the first comprehensive classification framework and architectural pattern atlas for ML-enabled systems; (2) identification of seven prevalent integration topologies and four model reuse patterns; and (3) uncovering critical interdisciplinary collaboration barriers in ML-software co-development. The findings bridge the methodological gap between data science and software engineering at the model embedding stage, providing industry-practical architectural guidelines that significantly enhance the maintainability and reusability of ML systems.
Small and medium-sized enterprises (SMEs) face significant challenges in operationalizing AI, primarily due to constrained resources, limited AI expertise, and the absence of lightweight, production-ready engineering and MLOps support. To address this gap, we propose the first lightweight AI engineering and MLOps blueprint framework specifically designed for SMEs. It integrates domain-customized reference architectures, automated toolchains, and iterative on-site validation mechanisms. Unlike generic enterprise-grade solutions, our blueprint prioritizes low entry barriers, high component reusability, and rapid deployment across the full AI lifecycle—encompassing model development, delivery, and operations. Empirical evaluation across multiple real-world business scenarios demonstrates an average 40% reduction in model delivery time and substantially improved development repeatability. Developer interviews confirm marked reductions in both technical adoption barriers and operational complexity. This work advances the scalable transfer of AI engineering practices from large enterprises to SMEs.
Addressing the acute shortage of AI/ML expertise in software engineering (SE), this study investigates the effectiveness and adoption barriers of AutoML for SE decision-making. Method: We systematically benchmark 12 state-of-the-art AutoML tools (e.g., H2O, Auto-sklearn, TPOT) on SE datasets and complement quantitative evaluation with surveys and expert interviews. Contribution/Results: Our empirical analysis reveals that AutoML-generated models achieve significantly higher average accuracy than manually tuned models on SE classification tasks. However, 83% of the tools lack automated feature engineering and deployment capabilities, and provide insufficient workflow support for non-ML experts—exposing a critical “pseudo end-to-end” limitation. The study identifies structural gaps in full-lifecycle automation and cross-role collaboration within current AutoML systems, thereby providing evidence-based insights and concrete design directions for next-generation AutoML tailored to SE contexts.
This study addresses the widespread absence of systematic instruction on building, testing, deploying, and maintaining AI/ML systems in current undergraduate software engineering (SE) curricula. It presents the first comprehensive delineation of core AI/ML topics essential for SE practice, integrating curriculum mapping analysis, instructor needs surveys, and structured modeling to identify critical content gaps in existing programs. Grounded in empirical evidence, the work proposes actionable pathways for embedding high-priority AI/ML themes into established SE courses. The resulting framework offers a practical, implementable guide to enhance SE education’s capacity to support the development of intelligent software systems.
This study addresses the prevalent ad hoc and non-standardized practices in model integration and deployment within MLOps projects, which often stem from a lack of systematic architectural guidance. To bridge this gap, the authors conduct a gray literature review of 103 online sources and apply thematic analysis to derive, for the first time, 25 architecturally significant best practices. These practices are systematically categorized into five thematic groups, with explicit articulation of each practice’s impact on overall system architecture. The resulting framework offers a structured, actionable set of guidelines for MLOps model integration and deployment, providing both researchers and engineering teams with a coherent theoretical foundation and practical reference for designing robust, scalable machine learning systems.
This work addresses the heavy reliance on human expertise in AI model development—spanning architecture design, representation engineering, and training optimization—and the limited scope of existing AutoML approaches. To bridge this gap, the authors propose AIBuildAI, a hierarchical multi-agent system powered by large language models that achieves end-to-end automation from natural language task descriptions to deployable models. The system comprises three coordinated agents responsible for design, coding, and tuning, leveraging multi-step reasoning and tool invocation. Evaluated on the MLE-Bench benchmark, AIBuildAI attains a medal rate of 63.1%, outperforming prior methods by a significant margin and matching the performance of experienced AI engineers.