Score
Designs, builds, and operates the systems and tooling required to run machine-learning model training workloads, including compute provisioning, distributed-training frameworks, job schedulers, data-loading and preprocessing pipelines, experiment tracking, and monitoring. Analyzes and optimizes training performance, cost, reliability, and reproducibility by implementing resource management, fault tolerance, automation, and deployment/CI practices for training pipelines.
Rapid growth in distributed DNN model size far outpaces hardware evolution, making it increasingly difficult to design training systems that simultaneously achieve high efficiency and sustainability. Method: This paper introduces the first three-dimensional evaluation framework—spanning workload abstraction, simulation infrastructure, and total cost of ownership (TCO)/carbon emission modeling—grounded in systematic literature review, multi-dimensional comparative modeling, and quantitative analysis. Contribution/Results: We identify common limitations of existing simulators in workload characterization, resource modeling, and environmental impact assessment; propose a structured capability comparison matrix to clarify assumptions, functional boundaries, and applicability of each tool; and expose cross-layer modeling gaps, distilling key open research challenges. Our framework provides a reproducible, extensible benchmark and decision-support foundation for co-designing efficient, low-carbon distributed training systems.
Medical AI deployment is hindered by insufficient production readiness of machine learning (ML) training pipelines. Method: This paper presents a progressive architectural evolution path—monolithic (chaotic) → modular monolithic → microservices—using SPIRA, a voice-based pre-diagnostic system for respiratory insufficiency, as a case study. It systematically introduces continuous training (CT) and a software-quality-attribute-driven MLOps governance framework tailored to healthcare, integrating modular design, microservice decomposition, and engineered CI/CD pipelines. Contribution/Results: The approach significantly improves pipeline maintainability, fault tolerance, and scalability, enabling stable, iterative evolution of SPIRA. It establishes an “agile ML + robust software engineering” co-design paradigm, delivering a reusable methodology and practical benchmark for engineering medical AI in highly regulated environments.
This work addresses the challenges of SLO violations and resource inefficiency in machine learning model serving caused by inadequate capacity planning. To this end, the authors propose an adaptive, feedback-driven load testing framework that formalizes the ML serving load testing process for the first time. The framework incorporates real-traffic-based workload calibration and a warm-up mechanism, combined with adaptive search, performance signal feedback control, convergence detection, and GPU monitoring to efficiently estimate the maximum sustainable throughput under SLO constraints. Evaluation across 14 industrial cases demonstrates that the approach reduces capacity estimation error from approximately 30% to 2–6%, with the warm-up mechanism improving accuracy by 22.2%. This significantly mitigates deployment incidents and enhances GPU resource utilization efficiency.
Reliability of large-scale ML training clusters is increasingly critical, yet the scaling behavior of failure impact remains poorly understood. This paper analyzes 11 months of operational data from two multi-tenant clusters—comprising 4 million ML training jobs and 150 million A100 GPU-hours—to establish, for the first time, a fine-grained failure taxonomy and reliability metric suite tailored to ML workloads. We propose the Effective Training Time Ratio (ETTR) modeling framework, enabling quantitative pre-evaluation of software-based fault-mitigation strategies. Through MTTF fitting and workload characterization, we reveal that although small tasks exhibit higher per-job robustness, their sheer volume makes them the dominant contributor to overall efficiency degradation. Our contributions include a scalable failure prediction model and a methodology for reliability assessment at scale—providing empirical foundations and theoretical insights for reliability-aware design of large-scale ML infrastructure.
Modern large-scale distributed training faces sharply diminishing returns in hardware scaling: as GPU counts reach thousands, communication overhead dominates performance bottlenecks, rendering conventional parallelism strategies—data, tensor, and pipeline parallelism—suboptimal. Method: Leveraging real-world LLM training workloads, this project establishes an empirical analytical framework spanning diverse model scales, hardware configurations, and parallelization strategies. It quantifies the nonlinear relationship between accelerator count and performance gain, precisely identifying critical inflection points across model, data, and compute scaling dimensions. Contribution/Results: We discover that low-communication “suboptimal” strategies become optimal at extreme scale; we empirically determine hardware selection criteria, cluster topology requirements, and optimal parallelism combinations for training billion-parameter models. Our findings provide actionable, deployment-ready optimization guidelines for trillion-parameter LLM training infrastructures.
In large-scale DNN distributed training, checkpointing is tightly coupled with model parallelism strategies and hardware topology, severely limiting fault tolerance and elastic scalability. To address this, we propose the “distributed storage, unified loading” paradigm: during saving, model parameters are stored in a distributed representation aligned with the current parallel configuration; during restoration, they are uniformly reconstructed into a logically consistent parameter view. We design a universal checkpoint format—incorporating merged parameter representations and mapping metadata—a Universal Checkpoint Language (UCL), and an on-demand state reconstruction mechanism, achieving, for the first time, full decoupling of checkpointing from parallel configurations. Evaluated on LLaMA, Bloom, and other mainstream large models under diverse parallelism paradigms—including tensor parallelism (TP), pipeline parallelism (PP), data parallelism (DP), and context parallelism (CP)—our approach reduces post-failure recovery time by 12–28% on average, significantly enhancing cross-configuration portability and system robustness.
To address the I/O bottleneck in ML training—causing low GPU utilization (often <50%)—this paper proposes a data-driven approach for I/O performance prediction and storage configuration optimization. We conduct systematic benchmarking across 141 configurations spanning diverse storage backends (NVMe SSDs, network-attached storage, in-memory filesystems), data formats, and access patterns. Leveraging key features—including batch size and throughput—we train an XGBoost regression model achieving an R² of 0.991 and a mean absolute error of only 11.8%. The model enables minute-scale recommendation of optimal storage configurations, accelerating configuration search by over an order of magnitude compared to empirical trial-and-error. Our core contribution is the first general-purpose, ML-training-pipeline-aware I/O performance prediction framework, which significantly improves GPU utilization and end-to-end training efficiency. All code and benchmark data are fully open-sourced, ensuring strong reproducibility and extensibility.
Existing large-scale model training systems struggle to flexibly compose diverse parallelization strategies, often relying on manual expert tuning and lacking generality. This work proposes a programmable distributed training system that enables users to declaratively specify composite parallelism strategies—such as data, pipeline, and expert parallelism—through model annotations and scheduling directives. These specifications are compiled via a unified intermediate representation (IR) into device-level execution plans, fully decoupling strategy definition from runtime execution over a global compute-communication DAG. The system is the first to support automatic compilation of user-defined composite strategies, matching the performance of established approaches like ZeRO while significantly improving both performance and memory efficiency in complex scenarios such as DeepSeek-V3’s DualPipe.
This work addresses customs clearance delays in global trade caused by ambiguous product descriptions and frequent updates to Harmonized System (HS) codes. To tackle this challenge, the authors propose a serverless MLOps framework that leverages event-driven pipelines and managed services to enable end-to-end, model-agnostic machine learning lifecycle management. The architecture supports automatic scaling, reproducible training, auditable deployment, and automated A/B testing, ensuring secure and seamless model transitions. By integrating custom text embeddings with models such as Text-CNN, the system achieves 98% accuracy on real-world HS code prediction tasks, meeting stringent service-level agreement (SLA) requirements. This approach significantly reduces long-term operational costs and establishes an efficient, cost-effective, and reproducible deployment paradigm for industrial-scale machine learning systems.
Traditional static data replication strategies in large-scale distributed systems struggle to adapt to dynamic workloads and sudden failures, resulting in low resource utilization and prolonged downtime. To address this, this paper proposes a machine learning–based adaptive replication mechanism that jointly optimizes failure prediction and replica placement using real-time monitoring data, integrating time-series forecasting with deep reinforcement learning. Its key innovation lies in the first end-to-end, fine-grained, online integration of predictive analytics and policy decision-making into a closed-loop self-evolving framework for replication strategy adaptation. Experimental evaluation demonstrates that the proposed approach reduces average downtime by 42.7% and improves storage resource utilization by 31.5% compared to representative static and heuristic baselines, while significantly enhancing fault tolerance resilience. These results validate the feasibility and effectiveness of intelligent, autonomous data replication in production-scale distributed systems.
This study systematically investigates the impact of network topology on collective communication performance in large-scale machine learning training. Focusing on Clos (fat-tree) and torus topologies, the work proposes an analytical model to quantitatively evaluate their completion times for operations such as AllReduce, while jointly accounting for network failures and task placement strategies. It establishes, for the first time, a quantitative relationship between topology structure and collective communication efficiency, demonstrating that Clos topologies consistently outperform torus networks across most scenarios—particularly in terms of communication latency, fault tolerance, and scheduling flexibility. These findings provide a rigorous theoretical foundation for designing network architectures tailored to ML training clusters.