hybrid parallelism tuning

Designs and evaluates hybrid parallelism configurations and memory-aware optimizer strategies for large-scale model training, including choosing pipeline and data-parallel degrees, microbatch sizes, and placement strategies to meet HBM and network capacity constraints. Builds methods to offload, shard, fuse, or otherwise modify optimizer state and update steps so as to reduce per-device memory footprint, preserve lossless parameter updates, and maximize training throughput under given resource limits.

hybridparallelismtuning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.35
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$191K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the computational and memory bottlenecks that hinder efficient scaling in large model training. To overcome the limitations of conventional point-wise optimizations, the authors propose a throughput-centric strategy that systematically integrates multiple techniques: optimized data loading (OVERLORD), CPU memory offloading (DeepSpeed ZeRO-Offload), distributed compilation (Triton-distributed), and hardware-level dynamic voltage and frequency scaling (DVFS). This holistic approach achieves a 4.5% improvement in end-to-end training throughput, substantially reduces training costs, and enables efficient training of models significantly larger than the memory capacity of a single GPU.

computational bottlenecklarge-scale AI systemsmemory bottleneck

This study systematically investigates hybrid parallelism strategies for large language models during both training and inference, aiming to balance computational, communication, and memory overheads. By constructing a mathematical cost model grounded in collective communication operations and integrating communication-computation overlap with automated strategy search, the work proposes a hybrid parallelism framework that achieves both efficiency and scalability. It is the first to unify theoretical modeling, automated search, and empirical evaluation across multiple hardware architectures, revealing the trade-offs among different parallelization strategies in training versus inference. The resulting framework provides reusable deployment guidelines for canonical model architectures, significantly enhancing distributed efficiency.

distributed parallelismhybrid parallelizationlarge language models

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training

Nov 20, 2024
JF
Jared Fernandez
🏛️ Meta | Carnegie Mellon University

Modern large-scale distributed training faces sharply diminishing returns in hardware scaling: as GPU counts reach thousands, communication overhead dominates performance bottlenecks, rendering conventional parallelism strategies—data, tensor, and pipeline parallelism—suboptimal. Method: Leveraging real-world LLM training workloads, this project establishes an empirical analytical framework spanning diverse model scales, hardware configurations, and parallelization strategies. It quantifies the nonlinear relationship between accelerator count and performance gain, precisely identifying critical inflection points across model, data, and compute scaling dimensions. Contribution/Results: We discover that low-communication “suboptimal” strategies become optimal at extreme scale; we empirically determine hardware selection criteria, cluster topology requirements, and optimal parallelism combinations for training billion-parameter models. Our findings provide actionable, deployment-ready optimization guidelines for trillion-parameter LLM training infrastructures.

Assessing diminishing returns in scaling accelerators for large model trainingEvaluating parallelization strategies to minimize distributed communication overheadOptimizing hardware configuration for efficient large-scale model training

Existing large-scale model training systems struggle to flexibly compose diverse parallelization strategies, often relying on manual expert tuning and lacking generality. This work proposes a programmable distributed training system that enables users to declaratively specify composite parallelism strategies—such as data, pipeline, and expert parallelism—through model annotations and scheduling directives. These specifications are compiled via a unified intermediate representation (IR) into device-level execution plans, fully decoupling strategy definition from runtime execution over a global compute-communication DAG. The system is the first to support automatic compilation of user-defined composite strategies, matching the performance of established approaches like ZeRO while significantly improving both performance and memory efficiency in complex scenarios such as DeepSeek-V3’s DualPipe.

distributed trainingflexibilitymodel parallelism

This work systematically investigates efficiency bottlenecks in large-scale LLM training across multi-GPU clusters (NVIDIA H100/H200, AMD MI250), focusing on the coupled effects of hardware utilization, power consumption, thermal throttling, and communication overhead. We conduct a multidimensional performance analysis of dense and sparse models using joint evaluation of tensor, pipeline, data, and expert parallelism—augmented with activation recomputation and compute-communication overlap. Key findings include: (i) scaling alone does not guarantee superior performance; smaller high-memory clusters outperform larger configurations in specific scenarios; (ii) tensor + pipeline parallelism often underutilizes interconnect bandwidth; and (iii) excessively large microbatches trigger power spikes and thermal throttling. Based on these insights, we propose parallelism strategy optimizations that jointly improve scalability and thermal stability. All experimental code is publicly released.

Analyzing power, performance, thermal impacts of parallelism strategiesCharacterizing LLM training efficiency across multi-GPU systemsEvaluating hardware utilization under different optimization techniques

Latest Papers

What's happening recently
View more

Training large language models (LLMs) on heterogeneous GPU clusters—including preemptible Spot instances—faces three key challenges: difficulty in coordinating asymmetric tensor and pipeline parallelism, inefficient gradient synchronization, and suboptimal memory-computation trade-offs. Method: This paper proposes AutoHet, the first system supporting fine-grained load allocation for asymmetric 3D parallelism (tensor, pipeline, and data parallelism). It formulates a joint optimization model to minimize per-step training time via device grouping and load balancing, and introduces a locality-aware, fast fault recovery mechanism tailored for Spot interruptions. Contribution/Results: Evaluated on three mixed-GPU cluster configurations (V100/A100/H100), AutoHet achieves up to 1.79× higher throughput than Megatron-LM and Whale when training three representative LLMs. Upon Spot instance preemption, its recovery speed is 4.38× faster than baseline approaches.

Addressing asymmetric pipeline parallelism and gradient synchronization challengesEnabling efficient recovery from spot instance preemptions in trainingOptimizing distributed training across heterogeneous GPU environments

Frontier models increasingly adopt Mixture-of-Experts (MoE) architectures to achieve large-model performance at reduced cost. However, training MoE models on HPC platforms is hindered by large memory footprints, frequent large-scale communication across heterogeneous networks, and severe workload imbalance. To characterize these challenges, we develop a mathematical model that quantifies memory, compute, and communication requirements for MoE configurations under various parallelization schemes, verified through micro-benchmarking, code instrumentation, and hardware profiling. Our analysis identifies performance bottlenecks: all-to-all latency at scale from expert parallelism, insufficient compute-communication overlap, low GPU utilization from imbalanced skinny GEMMs, and the absence of platform-aware hybrid parallelization strategies. To address these, we introduce Piper, a framework that leverages resource modeling to identify efficient training strategies for MoE models on target HPC platforms, applying pipeline parallelism with optimized schedules. Piper achieves 2-3.5X higher MFU than state-of-the-art frameworks such as X-MoE, and a novel all-to-all algorithm delivers 1.2-9X bandwidth over vendor implementation.

communication overheadHPClarge-scale training

This work addresses the substantial accelerator memory consumption of model parameters, gradients, and optimizer states in standard mixed-precision training, which hinders the scalability of large models. The authors propose a memory-efficient training method that significantly reduces quantization error in 8-bit optimizer states through compact master weight partitioning and a novel compression-expansion function. By integrating 16-bit gradients, an improved weight splitting strategy, and a gradient checkpointing mechanism, the approach remains compatible with mainstream optimizers such as SGD, AdamW, and Lion. The method reduces AdamW’s per-parameter memory footprint from 16 bytes to 7 bytes (or 5 bytes when gradients are released) and halves model checkpoint size, achieving lossless training quality across multiple vision and language benchmark tasks.

accelerator memorylarge language modelsmemory efficient training

This work addresses the memory bottlenecks and communication overheads encountered when training trillion-parameter Mixture-of-Experts (MoE) models with million-token context lengths. To overcome these challenges, the authors propose a “Mixture-of-Parallelisms” paradigm that synergistically integrates data, tensor, expert, and pipeline parallelism, complemented by memory-efficient optimizer state management and communication scheduling strategies. This approach enables, for the first time, lossless training of trillion-parameter MoE models at 1M-token context lengths while substantially reducing hardware requirements. Experimental results demonstrate that on a cluster of 12 nodes equipped with 8×H200 GPUs each, the method achieves per-GPU throughput 4.7–8.2× higher than the FSDP2 baseline, which suffers from out-of-memory errors even at context lengths of 64–128K tokens.

large-scale traininglong context lengthmemory efficiency

This work addresses the significant degradation in inference throughput caused by GPU memory constraints when concurrently deploying multiple large language models on shared heterogeneous hardware, where resource scheduling, model offloading, and preemption become critical bottlenecks. Through empirical methodologies—including cross-platform performance profiling, layer-wise offloading experiments, and fine-grained decomposition of preemption overhead—the study systematically uncovers, for the first time, the nonlinear relationship between offloading and throughput decline. It further identifies model state reloading as the primary source of preemption overhead. The findings reveal that smaller models are more sensitive to reduced GPU residency, and that such overhead is jointly influenced by model architecture and hardware characteristics. These insights motivate a scheduler design that integrates model-specific sensitivity with data migration costs, offering crucial guidance for building efficient multi-model serving systems.

CPU-GPU offloadingGPU memory constraintsheterogeneous hardware

Hot Scholars

MH

Mingyi Hong

Associate Professor, University of Minnesota; Amazon AGI
Machine LearningOptimizationGenerative AISignal processing
TL

Trung Le

Faculty of Information Technology, Monash University, Australia
Adversarial Machine LearningGenerative ModelsModel UnlearningModel Editing
MA

Maneesh Agrawala

Stanford University
GraphicsComputer GraphicsHCIVisualization
RP

Roberto Passerone

Professor of Electrical Engineering and Computer Sciences, University of Trento
Design methodologiessystem level designcontract-based designmodel-based design
LM

Lovish Madaan

AI at Meta & University College London
Machine LearningNatural Language Processing