Score
Designs, builds, and maintains software systems and engineering infrastructure for developing, deploying, and operating machine learning models, including model serving and inference pipelines. Implements and optimizes runtime performance (including accelerator- and hardware-specific integration), ensures reproducible engineering workflows, and develops the tooling for efficient model deployment, monitoring, and maintenance.
Real-world deployment of AI systems faces dual challenges: heterogeneous data floods and stringent low-latency requirements—straining conventional software architectures to their limits. To address this, we conduct the first systematic analysis of 217 production-deployed ML systems from a Data-Oriented Architecture (DOA) perspective, uncovering implicit DOA design patterns—particularly in loose coupling, decentralization, and data-driven orchestration. Leveraging systematic literature review, architectural pattern extraction, and requirement-to-design mapping analysis, we synthesize the first empirically grounded DOA practice guide and open-challenge taxonomy tailored to real-world ML deployments. Furthermore, we propose actionable, reusable recommendations for ML system deployment and introduce a novel architecture evaluation framework. Collectively, these contributions bridge the critical knowledge gap between DOA theory and industrial engineering practice.
This work addresses the challenges of SLO violations and resource inefficiency in machine learning model serving caused by inadequate capacity planning. To this end, the authors propose an adaptive, feedback-driven load testing framework that formalizes the ML serving load testing process for the first time. The framework incorporates real-traffic-based workload calibration and a warm-up mechanism, combined with adaptive search, performance signal feedback control, convergence detection, and GPU monitoring to efficiently estimate the maximum sustainable throughput under SLO constraints. Evaluation across 14 industrial cases demonstrates that the approach reduces capacity estimation error from approximately 30% to 2–6%, with the warm-up mechanism improving accuracy by 22.2%. This significantly mitigates deployment incidents and enhances GPU resource utilization efficiency.
To address low testing and debugging efficiency and immature toolchains when deploying rapidly evolving large language models (LLMs) on emerging platforms (e.g., browsers, mobile devices), this paper proposes TapML—a top-down, test-driven framework. Methodologically, TapML introduces (1) the first operator-level test pruning technique that automatically generates high-coverage, realistic test inputs; (2) a progressive cross-platform migration strategy that significantly narrows the scope for compound error localization; and (3) native backend support for Metal and WebGPU, with deep integration into MLC-LLM. Evaluated over two years, TapML has enabled efficient deployment of 105 emerging models—spanning 27 distinct architectures—across five platform categories, reducing average deployment time by 42%. It has since become the default development paradigm for MLC-LLM.
Addressing the longstanding challenge in software engineering (SE) of balancing quality assurance with development efficiency, this study proposes a systematically derived and optimized machine learning (ML) pipeline framework tailored for SE tasks. Methodologically, the pipeline integrates automated data acquisition, SMOTE-based class balancing, SZZ-inspired feature selection, ensemble models (Random Forest and Gradient Boosting), and a novel evaluation metric—Balanced Accuracy Metric (BAM)—alongside standard metrics (AUC, F1-score, precision) and bootstrap resampling for robust validation. Key contributions include: (1) the first structured, end-to-end optimization framework for SE-specific ML pipelines; (2) empirical evidence demonstrating superior defect prediction performance of ensemble methods over individual classifiers; and (3) identification of two critical research gaps—insufficient data standardization and lack of model interpretability—thereby establishing foundational theoretical insights and actionable directions for intelligent SE research.
This study addresses the prevailing gap in AI education, which emphasizes model development while neglecting system engineering practices, leaving students ill-equipped to handle real-world challenges such as architectural design, deployment, and monitoring. To bridge this gap, the authors implemented a master’s-level course in which students built a movie recommendation system under realistic constraints, with a focus on integrating AI components into robust software systems, adopting data-driven machine learning practices, and cultivating systems-level thinking. Using a mixed-methods approach—combining analysis of student project artifacts with survey data—the research evaluates learners’ performance in architectural decision-making, integration of heterogeneous models, and adaptation to evolving requirements. Findings reveal common difficulties students encounter in AI system engineering and demonstrate the course’s effectiveness in addressing critical deficiencies in AI engineering education and enhancing systems-aware competencies.
This study addresses the challenges of integrating machine learning (ML) models into software systems—namely, poor integration practices, low reusability, and unclear architectural boundaries. It presents the first large-scale empirical investigation across 2,928 open-source ML-enabled systems. Leveraging GitHub code mining, static analysis, topic modeling, and architectural pattern identification, the work systematically characterizes ML integration topologies, code/model reuse practices, and maintenance bottlenecks. Key contributions include: (1) the first comprehensive classification framework and architectural pattern atlas for ML-enabled systems; (2) identification of seven prevalent integration topologies and four model reuse patterns; and (3) uncovering critical interdisciplinary collaboration barriers in ML-software co-development. The findings bridge the methodological gap between data science and software engineering at the model embedding stage, providing industry-practical architectural guidelines that significantly enhance the maintainability and reusability of ML systems.
Existing tools struggle to efficiently and accurately perform full-stack architectural analysis of machine learning infrastructure spanning from microwatt-scale devices to gigawatt-scale data centers. This work proposes a first-principles-based analytical modeling framework that decouples computational demand from hardware supply and environmental context through a “demand–supply” abstraction. The framework introduces a “walls-of-systems” taxonomy, a dimensionally rigorous Python engine, 22 classes of system constraints, 28 composable solvers, and a typed input registry with provenance tracking to ensure unit consistency and traceability. It enables sub-second design space exploration, precisely identifies system bottlenecks, and automatically generates optimal hardware configurations covering the entire machine learning lifecycle.
This study addresses the lack of systematic empirical investigation into the deployment of large language model (LLM) inference frameworks in real-world software systems. Through large-scale analysis of open-source projects, combining static code analysis with metadata mining, this work systematically characterizes the adoption patterns and integration strategies of prominent frameworks—including vLLM, SGLang, TensorRT-LLM, LMDeploy, and FlashInfer—and examines their relationships with model type, scale, modality, and deployment environment. The findings reveal that vLLM exhibits the highest adoption rate, while multi-framework co-deployment remains limited yet holds complementary potential. Framework selection is strongly driven by model characteristics and application scenarios, effectively supporting diverse system designs such as reinforcement learning inference, multimodal generation, and microservice architectures.
This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.