Score
Designs and evaluates hybrid parallelism configurations and memory-aware optimizer strategies for large-scale model training, including choosing pipeline and data-parallel degrees, microbatch sizes, and placement strategies to meet HBM and network capacity constraints. Builds methods to offload, shard, fuse, or otherwise modify optimizer state and update steps so as to reduce per-device memory footprint, preserve lossless parameter updates, and maximize training throughput under given resource limits.
This work addresses the common practice of optimizing large language model training in isolation due to constraints in data, memory, and compute resources, which lacks a holistic perspective. The authors propose a unified resource-aware decision framework that systematically integrates data efficiency, memory compression, and compute budget-aware strategies. They demonstrate that optimal data selection critically depends on both task objectives and resource constraints, and reveal that memory—not computational power—is often the primary bottleneck in fine-tuning. By combining learning-dynamics-, gradient-, and influence-based data pruning with adaptive stopping criteria and inference allocation mechanisms, the framework establishes a systematic approach for efficient training and deployment under limited resources, substantially improving overall resource utilization efficiency.
This work addresses the computational and memory bottlenecks that hinder efficient scaling in large model training. To overcome the limitations of conventional point-wise optimizations, the authors propose a throughput-centric strategy that systematically integrates multiple techniques: optimized data loading (OVERLORD), CPU memory offloading (DeepSpeed ZeRO-Offload), distributed compilation (Triton-distributed), and hardware-level dynamic voltage and frequency scaling (DVFS). This holistic approach achieves a 4.5% improvement in end-to-end training throughput, substantially reduces training costs, and enables efficient training of models significantly larger than the memory capacity of a single GPU.
This study systematically investigates hybrid parallelism strategies for large language models during both training and inference, aiming to balance computational, communication, and memory overheads. By constructing a mathematical cost model grounded in collective communication operations and integrating communication-computation overlap with automated strategy search, the work proposes a hybrid parallelism framework that achieves both efficiency and scalability. It is the first to unify theoretical modeling, automated search, and empirical evaluation across multiple hardware architectures, revealing the trade-offs among different parallelization strategies in training versus inference. The resulting framework provides reusable deployment guidelines for canonical model architectures, significantly enhancing distributed efficiency.
Modern large-scale distributed training faces sharply diminishing returns in hardware scaling: as GPU counts reach thousands, communication overhead dominates performance bottlenecks, rendering conventional parallelism strategies—data, tensor, and pipeline parallelism—suboptimal. Method: Leveraging real-world LLM training workloads, this project establishes an empirical analytical framework spanning diverse model scales, hardware configurations, and parallelization strategies. It quantifies the nonlinear relationship between accelerator count and performance gain, precisely identifying critical inflection points across model, data, and compute scaling dimensions. Contribution/Results: We discover that low-communication “suboptimal” strategies become optimal at extreme scale; we empirically determine hardware selection criteria, cluster topology requirements, and optimal parallelism combinations for training billion-parameter models. Our findings provide actionable, deployment-ready optimization guidelines for trillion-parameter LLM training infrastructures.
Existing large-scale model training systems struggle to flexibly compose diverse parallelization strategies, often relying on manual expert tuning and lacking generality. This work proposes a programmable distributed training system that enables users to declaratively specify composite parallelism strategies—such as data, pipeline, and expert parallelism—through model annotations and scheduling directives. These specifications are compiled via a unified intermediate representation (IR) into device-level execution plans, fully decoupling strategy definition from runtime execution over a global compute-communication DAG. The system is the first to support automatic compilation of user-defined composite strategies, matching the performance of established approaches like ZeRO while significantly improving both performance and memory efficiency in complex scenarios such as DeepSeek-V3’s DualPipe.
This work systematically investigates efficiency bottlenecks in large-scale LLM training across multi-GPU clusters (NVIDIA H100/H200, AMD MI250), focusing on the coupled effects of hardware utilization, power consumption, thermal throttling, and communication overhead. We conduct a multidimensional performance analysis of dense and sparse models using joint evaluation of tensor, pipeline, data, and expert parallelism—augmented with activation recomputation and compute-communication overlap. Key findings include: (i) scaling alone does not guarantee superior performance; smaller high-memory clusters outperform larger configurations in specific scenarios; (ii) tensor + pipeline parallelism often underutilizes interconnect bandwidth; and (iii) excessively large microbatches trigger power spikes and thermal throttling. Based on these insights, we propose parallelism strategy optimizations that jointly improve scalability and thermal stability. All experimental code is publicly released.
Training large language models (LLMs) on heterogeneous GPU clusters—including preemptible Spot instances—faces three key challenges: difficulty in coordinating asymmetric tensor and pipeline parallelism, inefficient gradient synchronization, and suboptimal memory-computation trade-offs. Method: This paper proposes AutoHet, the first system supporting fine-grained load allocation for asymmetric 3D parallelism (tensor, pipeline, and data parallelism). It formulates a joint optimization model to minimize per-step training time via device grouping and load balancing, and introduces a locality-aware, fast fault recovery mechanism tailored for Spot interruptions. Contribution/Results: Evaluated on three mixed-GPU cluster configurations (V100/A100/H100), AutoHet achieves up to 1.79× higher throughput than Megatron-LM and Whale when training three representative LLMs. Upon Spot instance preemption, its recovery speed is 4.38× faster than baseline approaches.
Frontier models increasingly adopt Mixture-of-Experts (MoE) architectures to achieve large-model performance at reduced cost. However, training MoE models on HPC platforms is hindered by large memory footprints, frequent large-scale communication across heterogeneous networks, and severe workload imbalance. To characterize these challenges, we develop a mathematical model that quantifies memory, compute, and communication requirements for MoE configurations under various parallelization schemes, verified through micro-benchmarking, code instrumentation, and hardware profiling. Our analysis identifies performance bottlenecks: all-to-all latency at scale from expert parallelism, insufficient compute-communication overlap, low GPU utilization from imbalanced skinny GEMMs, and the absence of platform-aware hybrid parallelization strategies. To address these, we introduce Piper, a framework that leverages resource modeling to identify efficient training strategies for MoE models on target HPC platforms, applying pipeline parallelism with optimized schedules. Piper achieves 2-3.5X higher MFU than state-of-the-art frameworks such as X-MoE, and a novel all-to-all algorithm delivers 1.2-9X bandwidth over vendor implementation.
This work addresses the substantial accelerator memory consumption of model parameters, gradients, and optimizer states in standard mixed-precision training, which hinders the scalability of large models. The authors propose a memory-efficient training method that significantly reduces quantization error in 8-bit optimizer states through compact master weight partitioning and a novel compression-expansion function. By integrating 16-bit gradients, an improved weight splitting strategy, and a gradient checkpointing mechanism, the approach remains compatible with mainstream optimizers such as SGD, AdamW, and Lion. The method reduces AdamW’s per-parameter memory footprint from 16 bytes to 7 bytes (or 5 bytes when gradients are released) and halves model checkpoint size, achieving lossless training quality across multiple vision and language benchmark tasks.
This work addresses the memory bottlenecks and communication overheads encountered when training trillion-parameter Mixture-of-Experts (MoE) models with million-token context lengths. To overcome these challenges, the authors propose a “Mixture-of-Parallelisms” paradigm that synergistically integrates data, tensor, expert, and pipeline parallelism, complemented by memory-efficient optimizer state management and communication scheduling strategies. This approach enables, for the first time, lossless training of trillion-parameter MoE models at 1M-token context lengths while substantially reducing hardware requirements. Experimental results demonstrate that on a cluster of 12 nodes equipped with 8×H200 GPUs each, the method achieves per-GPU throughput 4.7–8.2× higher than the FSDP2 baseline, which suffers from out-of-memory errors even at context lengths of 64–128K tokens.
This work addresses the significant degradation in inference throughput caused by GPU memory constraints when concurrently deploying multiple large language models on shared heterogeneous hardware, where resource scheduling, model offloading, and preemption become critical bottlenecks. Through empirical methodologies—including cross-platform performance profiling, layer-wise offloading experiments, and fine-grained decomposition of preemption overhead—the study systematically uncovers, for the first time, the nonlinear relationship between offloading and throughput decline. It further identifies model state reloading as the primary source of preemption overhead. The findings reveal that smaller models are more sensitive to reduced GPU residency, and that such overhead is jointly influenced by model architecture and hardware characteristics. These insights motivate a scheduler design that integrates model-specific sensitivity with data migration costs, offering crucial guidance for building efficient multi-model serving systems.