distributed training

Designs and implements training systems and pipelines that partition models and data across multiple devices or sites (e.g., GPUs and servers), orchestrate parallel optimization and communication to scale to large datasets and resolutions, and support iterative or online loops such as reinforcement learning. Builds fault-tolerant, memory- and resource-efficient methods — including model/data sharding, gradient accumulation, checkpointing and recovery, synchronization-reduction and topology-adaptive scheduling — and composes multi-stage and multi-task training workflows that operate under constrained, heterogeneous, or WAN conditions.

distributedtraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.73
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$219K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing large-scale model training systems struggle to flexibly compose diverse parallelization strategies, often relying on manual expert tuning and lacking generality. This work proposes a programmable distributed training system that enables users to declaratively specify composite parallelism strategies—such as data, pipeline, and expert parallelism—through model annotations and scheduling directives. These specifications are compiled via a unified intermediate representation (IR) into device-level execution plans, fully decoupling strategy definition from runtime execution over a global compute-communication DAG. The system is the first to support automatic compilation of user-defined composite strategies, matching the performance of established approaches like ZeRO while significantly improving both performance and memory efficiency in complex scenarios such as DeepSeek-V3’s DualPipe.

distributed trainingflexibilitymodel parallelism

This work addresses the inefficiency of conventional synchronous training in large-scale GPU clusters—specifically those with over 100,000 GPUs—where frequent hardware failures and prolonged recovery procedures severely degrade performance. The authors propose a fault-tolerant hybrid sharded data parallelism (FT-HSDP) paradigm that, for the first time, treats data-parallel replicas as independent fault-tolerance units. By integrating dynamic participant management and a non-blocking catch-up mechanism, FT-HSDP enables localized fault recovery: only affected replicas are restarted while others continue training uninterrupted. The system employs a CPU-coordinated, GPU-executed fault-tolerant all-reduce protocol (FTAR) that supports efficient asynchronous recovery without compromising model accuracy. Experiments on a 100,000-GPU cluster demonstrate that this approach reduces fault-induced training stalls from 10 minutes to 3 minutes and increases effective training time from 44% to 80%.

fault toleranceGPU failureslarge-scale training

Universal Checkpointing: Efficient and Flexible Checkpointing for Large Scale Distributed Training

Jun 27, 2024
XL
Xinyu Lian
🏛️ University of Illinois at Urbana-Champaign | Microsoft | StasoSphere

In large-scale DNN distributed training, checkpointing is tightly coupled with model parallelism strategies and hardware topology, severely limiting fault tolerance and elastic scalability. To address this, we propose the “distributed storage, unified loading” paradigm: during saving, model parameters are stored in a distributed representation aligned with the current parallel configuration; during restoration, they are uniformly reconstructed into a logically consistent parameter view. We design a universal checkpoint format—incorporating merged parameter representations and mapping metadata—a Universal Checkpoint Language (UCL), and an on-demand state reconstruction mechanism, achieving, for the first time, full decoupling of checkpointing from parallel configurations. Evaluated on LLaMA, Bloom, and other mainstream large models under diverse parallelism paradigms—including tensor parallelism (TP), pipeline parallelism (PP), data parallelism (DP), and context parallelism (CP)—our approach reduces post-failure recovery time by 12–28% on average, significantly enhancing cross-configuration portability and system robustness.

Decouples checkpoint structure from hardware configurationsEnables reconfigurable parallelism in large-scale DNN trainingSupports flexible mapping of checkpoint state to parallelism strategies

Optimizing Distributed Training Approaches for Scaling Neural Networks

Mar 29, 2025
VB
V. Baligodugula
🏛️ Wright State University

Static parallelization strategies in large-scale distributed neural network training suffer from poor resource adaptability, leading to efficiency bottlenecks. To address this, we systematically evaluate the performance boundaries of data parallelism, model parallelism, and hybrid parallelism, and propose a dynamic, topology- and resource-aware scheduling algorithm. This algorithm enables online switching of parallelization strategies within a hybrid parallel framework during training, jointly optimizing communication overhead, computational load balance, and memory constraints. On the CIFAR-100 image classification benchmark, hybrid parallelism achieves a 3.2× speedup over single-GPU training with no accuracy degradation; integrating our adaptive scheduler further improves end-to-end training efficiency by 18%. To the best of our knowledge, this work is the first to incorporate dynamic strategy switching into the hybrid parallel training pipeline, establishing a scalable new paradigm for efficient large-model training in heterogeneous resource environments.

Compare distributed training strategies for large neural networksEvaluate performance on CIFAR-100 image classification tasksPropose adaptive algorithm to optimize training efficiency

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training

Nov 20, 2024
JF
Jared Fernandez
🏛️ Meta | Carnegie Mellon University

Modern large-scale distributed training faces sharply diminishing returns in hardware scaling: as GPU counts reach thousands, communication overhead dominates performance bottlenecks, rendering conventional parallelism strategies—data, tensor, and pipeline parallelism—suboptimal. Method: Leveraging real-world LLM training workloads, this project establishes an empirical analytical framework spanning diverse model scales, hardware configurations, and parallelization strategies. It quantifies the nonlinear relationship between accelerator count and performance gain, precisely identifying critical inflection points across model, data, and compute scaling dimensions. Contribution/Results: We discover that low-communication “suboptimal” strategies become optimal at extreme scale; we empirically determine hardware selection criteria, cluster topology requirements, and optimal parallelism combinations for training billion-parameter models. Our findings provide actionable, deployment-ready optimization guidelines for trillion-parameter LLM training infrastructures.

Assessing diminishing returns in scaling accelerators for large model trainingEvaluating parallelization strategies to minimize distributed communication overheadOptimizing hardware configuration for efficient large-scale model training

Latest Papers

What's happening recently
View more

This work addresses the challenges of low communication efficiency and system complexity in large-scale AI training across geographically distributed data centers. It presents the first systematic characterization and joint optimization of three critical dimensions in “scale-across” training: parallelism strategy deployment, job scheduling, and network transport. By co-designing parallel placement, scheduling policies, and advanced networking techniques—and validating the approach through both real-world testbeds and large-scale simulations—the study achieves end-to-end global optimization of computation and communication. Experimental results demonstrate that the proposed solution improves training throughput by up to 64.62% over current production configurations and by 37.59% compared to state-of-the-art baselines.

communication optimizationdata centerdistributed training

This work addresses the longstanding limitation of large language model (LLM) training, which has been confined to high-performance data centers and unable to leverage the vast pool of heterogeneous, unreliable consumer-grade GPUs available across the internet. To overcome this, the authors introduce Agora, a novel system that pioneers a “protocol learning” paradigm, enabling permissionless, decentralized collective pretraining through communication-efficient pipeline parallelism, asynchronous optimization, and robust fault tolerance. In this framework, no participant holds the full model; instead, each retains only a shard of the parameters, ensuring collective ownership and inherent openness. The system successfully trained Pluralis-8B (8.6 billion parameters) on 500 billion FineWeb-Edu tokens using 330 dynamically joining and leaving consumer GPUs over 40 days, achieving 63% of the computational efficiency of a centralized baseline while matching its convergence performance closely.

distributed trainingheterogeneous GPUsinternet-scale compute

Despite growing interest in Mixture-of-Experts (MoE) models, large-scale end-to-end pretraining of MoE architectures on pure AMD hardware—specifically the MI300X GPU and Pollara interconnect—remains unexplored, raising questions about the platform’s readiness for state-of-the-art LLM training. Method: We introduce MI300X-aware Transformer module sizing guidelines, conduct full-stack Pollara communication microbenchmarks, and integrate optimized All-reduce/Reduce-scatter primitives, fault-tolerant training, and checkpoint reshaping. Contribution/Results: We successfully complete end-to-end pretraining of the ZAYA1-base MoE model (760M activated parameters, 8.3B total parameters) on native AMD infrastructure. Evaluated on reasoning, mathematics, and code generation, ZAYA1-base substantially outperforms Llama-3-8B and OLMoE, matching the performance of Qwen3-4B and Gemma3-12B. This demonstrates, for the first time, that AMD’s hardware-software stack achieves full maturity for high-performance MoE model training.

Develops AMD hardware optimization for large-scale mixture-of-experts foundation model trainingIntroduces transformer sizing rules optimizing training throughput and inference latencyProvides system design guidance through cluster networking characterization and microbenchmarks

Frontier models increasingly adopt Mixture-of-Experts (MoE) architectures to achieve large-model performance at reduced cost. However, training MoE models on HPC platforms is hindered by large memory footprints, frequent large-scale communication across heterogeneous networks, and severe workload imbalance. To characterize these challenges, we develop a mathematical model that quantifies memory, compute, and communication requirements for MoE configurations under various parallelization schemes, verified through micro-benchmarking, code instrumentation, and hardware profiling. Our analysis identifies performance bottlenecks: all-to-all latency at scale from expert parallelism, insufficient compute-communication overlap, low GPU utilization from imbalanced skinny GEMMs, and the absence of platform-aware hybrid parallelization strategies. To address these, we introduce Piper, a framework that leverages resource modeling to identify efficient training strategies for MoE models on target HPC platforms, applying pipeline parallelism with optimized schedules. Piper achieves 2-3.5X higher MFU than state-of-the-art frameworks such as X-MoE, and a novel all-to-all algorithm delivers 1.2-9X bandwidth over vendor implementation.

communication overheadHPClarge-scale training

This work addresses the orchestration bottlenecks faced by ultra-large-scale Sim-AI workflows on leadership-class supercomputers, which arise from task heterogeneity and extreme ensemble sizes. To overcome these challenges, the authors propose EnsembleLauncher, a recursively hierarchical and fully decentralized workflow orchestrator that introduces a decentralized control plane and a programmable scheduling policy interface, thereby surpassing conventional tools in both scalability and scheduling flexibility. Experiments on the Aurora supercomputer demonstrate that EnsembleLauncher can efficiently schedule system-wide resources to support up to 8 million serial tasks, achieving more than a fourfold performance improvement over state-of-the-art alternatives. Furthermore, it significantly enhances resource utilization for workloads with high task variance and active learning pipelines.

exascaleorchestration bottlenecksscalability

Hot Scholars

FF

Fangcheng Fu

Shanghai Jiao Tong University
machine learningdeep learningMLSysdistributed computation
EB

Eugene Belilovsky

Associate Professor, Concordia University and Mila Quebec AI Institute
Distributed LearningContinual LearningFederated LearningLearned Optimizers
TH

Torsten Hoefler

Professor of Computer Science at ETH Zurich
High Performance ComputingDeep LearningNetworkingMessage Passing Interface
RY

Ran Yan

University of California, Los Angeles
Medical Imaging