Score
Designs and implements systems, pipelines, and procedures for training machine learning models at scale, including distributed/data/model parallelism, batching, checkpointing, mixed-precision and gradient-accumulation, efficient communication, resource scheduling, and dataset sharding. Builds and analyzes scalability, throughput, convergence behavior, cost efficiency, and fault-tolerance of large-scale model training workflows.
Modern large-scale distributed training faces sharply diminishing returns in hardware scaling: as GPU counts reach thousands, communication overhead dominates performance bottlenecks, rendering conventional parallelism strategies—data, tensor, and pipeline parallelism—suboptimal. Method: Leveraging real-world LLM training workloads, this project establishes an empirical analytical framework spanning diverse model scales, hardware configurations, and parallelization strategies. It quantifies the nonlinear relationship between accelerator count and performance gain, precisely identifying critical inflection points across model, data, and compute scaling dimensions. Contribution/Results: We discover that low-communication “suboptimal” strategies become optimal at extreme scale; we empirically determine hardware selection criteria, cluster topology requirements, and optimal parallelism combinations for training billion-parameter models. Our findings provide actionable, deployment-ready optimization guidelines for trillion-parameter LLM training infrastructures.
Existing large-scale model training systems struggle to flexibly compose diverse parallelization strategies, often relying on manual expert tuning and lacking generality. This work proposes a programmable distributed training system that enables users to declaratively specify composite parallelism strategies—such as data, pipeline, and expert parallelism—through model annotations and scheduling directives. These specifications are compiled via a unified intermediate representation (IR) into device-level execution plans, fully decoupling strategy definition from runtime execution over a global compute-communication DAG. The system is the first to support automatic compilation of user-defined composite strategies, matching the performance of established approaches like ZeRO while significantly improving both performance and memory efficiency in complex scenarios such as DeepSeek-V3’s DualPipe.
Static parallelization strategies in large-scale distributed neural network training suffer from poor resource adaptability, leading to efficiency bottlenecks. To address this, we systematically evaluate the performance boundaries of data parallelism, model parallelism, and hybrid parallelism, and propose a dynamic, topology- and resource-aware scheduling algorithm. This algorithm enables online switching of parallelization strategies within a hybrid parallel framework during training, jointly optimizing communication overhead, computational load balance, and memory constraints. On the CIFAR-100 image classification benchmark, hybrid parallelism achieves a 3.2× speedup over single-GPU training with no accuracy degradation; integrating our adaptive scheduler further improves end-to-end training efficiency by 18%. To the best of our knowledge, this work is the first to incorporate dynamic strategy switching into the hybrid parallel training pipeline, establishing a scalable new paradigm for efficient large-model training in heterogeneous resource environments.
In large-scale DNN distributed training, checkpointing is tightly coupled with model parallelism strategies and hardware topology, severely limiting fault tolerance and elastic scalability. To address this, we propose the “distributed storage, unified loading” paradigm: during saving, model parameters are stored in a distributed representation aligned with the current parallel configuration; during restoration, they are uniformly reconstructed into a logically consistent parameter view. We design a universal checkpoint format—incorporating merged parameter representations and mapping metadata—a Universal Checkpoint Language (UCL), and an on-demand state reconstruction mechanism, achieving, for the first time, full decoupling of checkpointing from parallel configurations. Evaluated on LLaMA, Bloom, and other mainstream large models under diverse parallelism paradigms—including tensor parallelism (TP), pipeline parallelism (PP), data parallelism (DP), and context parallelism (CP)—our approach reduces post-failure recovery time by 12–28% on average, significantly enhancing cross-configuration portability and system robustness.
This work systematically investigates efficiency bottlenecks in large-scale LLM training across multi-GPU clusters (NVIDIA H100/H200, AMD MI250), focusing on the coupled effects of hardware utilization, power consumption, thermal throttling, and communication overhead. We conduct a multidimensional performance analysis of dense and sparse models using joint evaluation of tensor, pipeline, data, and expert parallelism—augmented with activation recomputation and compute-communication overlap. Key findings include: (i) scaling alone does not guarantee superior performance; smaller high-memory clusters outperform larger configurations in specific scenarios; (ii) tensor + pipeline parallelism often underutilizes interconnect bandwidth; and (iii) excessively large microbatches trigger power spikes and thermal throttling. Based on these insights, we propose parallelism strategy optimizations that jointly improve scalability and thermal stability. All experimental code is publicly released.
This work addresses the challenges of low communication efficiency and system complexity in large-scale AI training across geographically distributed data centers. It presents the first systematic characterization and joint optimization of three critical dimensions in “scale-across” training: parallelism strategy deployment, job scheduling, and network transport. By co-designing parallel placement, scheduling policies, and advanced networking techniques—and validating the approach through both real-world testbeds and large-scale simulations—the study achieves end-to-end global optimization of computation and communication. Experimental results demonstrate that the proposed solution improves training throughput by up to 64.62% over current production configurations and by 37.59% compared to state-of-the-art baselines.
To address the lack of unified benchmarks for model performance evaluation on high-dimensional big data in both local and distributed environments, this work designs an end-to-end evaluation framework covering three representative tasks—Epsilon (numerical regression), RestMex (text classification), and IMDb (movie feature analysis). Leveraging Apache Spark (Scala), we establish a reproducible heterogeneous computing experimental infrastructure to systematically compare traditional machine learning and deep learning models across accuracy, training efficiency, and resource consumption. This study presents the first pedagogically implemented standardized benchmark supporting multiple models, multimodal data, and diverse deployment scenarios, empirically uncovering performance bottlenecks and architectural trade-offs inherent in distributed scaling. The outcomes include an open-source evaluation pipeline, a standardized reporting template, and a reusable teaching paradigm—providing empirical foundations for AI system selection and optimization in big data contexts.
Despite growing interest in Mixture-of-Experts (MoE) models, large-scale end-to-end pretraining of MoE architectures on pure AMD hardware—specifically the MI300X GPU and Pollara interconnect—remains unexplored, raising questions about the platform’s readiness for state-of-the-art LLM training. Method: We introduce MI300X-aware Transformer module sizing guidelines, conduct full-stack Pollara communication microbenchmarks, and integrate optimized All-reduce/Reduce-scatter primitives, fault-tolerant training, and checkpoint reshaping. Contribution/Results: We successfully complete end-to-end pretraining of the ZAYA1-base MoE model (760M activated parameters, 8.3B total parameters) on native AMD infrastructure. Evaluated on reasoning, mathematics, and code generation, ZAYA1-base substantially outperforms Llama-3-8B and OLMoE, matching the performance of Qwen3-4B and Gemma3-12B. This demonstrates, for the first time, that AMD’s hardware-software stack achieves full maturity for high-performance MoE model training.
This work addresses the challenges of SLO violations and resource inefficiency in machine learning model serving caused by inadequate capacity planning. To this end, the authors propose an adaptive, feedback-driven load testing framework that formalizes the ML serving load testing process for the first time. The framework incorporates real-traffic-based workload calibration and a warm-up mechanism, combined with adaptive search, performance signal feedback control, convergence detection, and GPU monitoring to efficiently estimate the maximum sustainable throughput under SLO constraints. Evaluation across 14 industrial cases demonstrates that the approach reduces capacity estimation error from approximately 30% to 2–6%, with the warm-up mechanism improving accuracy by 22.2%. This significantly mitigates deployment incidents and enhances GPU resource utilization efficiency.
This study systematically investigates hybrid parallelism strategies for large language models during both training and inference, aiming to balance computational, communication, and memory overheads. By constructing a mathematical cost model grounded in collective communication operations and integrating communication-computation overlap with automated strategy search, the work proposes a hybrid parallelism framework that achieves both efficiency and scalability. It is the first to unify theoretical modeling, automated search, and empirical evaluation across multiple hardware architectures, revealing the trade-offs among different parallelization strategies in training versus inference. The resulting framework provides reusable deployment guidelines for canonical model architectures, significantly enhancing distributed efficiency.