efficient pretraining

Designs and implements asynchronous pipeline-parallel pretraining systems and schedules that minimize pipeline bubbles and control gradient delay (e.g., bubble-free, pipedream-2bw, PaCI), including building the scheduler/protocol, integration with GPU/TPU execution, and stability mitigations. Measures and optimizes wall-clock throughput and training time, trades off compute efficiency versus model accuracy, and produces reproducible, stable pretraining routines and reports.

efficientpretraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.33
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$183K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the inefficiencies in pipeline parallel training, where synchronous methods suffer from bubble-induced idle time and asynchronous approaches introduce weight version inconsistency. The authors propose PACI, a novel method that explicitly controls the evolution rate of parameter versions through local gradient accumulation. PACI enables bubble-free asynchronous pipeline execution without requiring weight stashing, prediction, or global synchronization, while strictly bounding weight drift between forward and backward passes. Notably, it achieves high throughput and training stability without additional memory overhead or extra parameter copies. In pretraining GPT-style models, PACI matches the final perplexity and peak memory usage of synchronous 1F1B-flush, attains full pipeline utilization, and reduces training time to target accuracy by up to 1.69×.

asynchronous trainingbubblesoptimization consistency

Asynchronous pipeline parallel training suffers from optimization instability due to gradient staleness, limiting its applicability in large-scale pretraining. This work systematically demonstrates the critical role of optimizers in mitigating the adverse effects of first-order gradient staleness and proposes an enhanced approach that integrates the MuOn optimizer with an optimizer-agnostic error feedback mechanism. The method is accompanied by theoretical convergence guarantees. Experimental results on models up to 10 billion parameters show that the proposed technique substantially narrows the performance gap with synchronous training, enabling stable and efficient large-scale asynchronous pretraining.

asynchronous pipeline parallelismgradient stalenesslarge-scale LLM pretraining

This work addresses the lack of non-convex convergence guarantees in PipeDream-style pipeline parallelism by proposing the Randomized PipeDream (RPD) framework, for which it establishes the first rigorous non-convex convergence theory. By introducing a randomized block SGD abstraction coupled with explicit modeling of communication delays, the analysis reveals that under steady-state conditions, the delay grows quadratically with the number of pipeline stages \(S\), leading to stale gradient terms scaling as \(\Theta(S^4)\). Empirical evaluations demonstrate that RPD outperforms LocalSGD in quadratic optimization and small-scale language model training, whereas LocalSGD exhibits superior performance as \(S\) increases in logistic regression tasks, highlighting a nuanced trade-off between the two methods across different problem settings.

convergencedistributed trainingmodel parallelism

PipeOptim: Ensuring Effective 1F1B Schedule with Optimizer-Dependent Weight Prediction

Dec 01, 2023
LG
Lei Guan
🏛️ National University of Defense Technology | Shanxi University

In asynchronous pipeline parallelism, the 1F1B (one-forward-one-backward) scheduling incurs weight inconsistency and staleness due to interleaved mini-batch execution across GPUs. To address this, we propose an optimizer-aware forward-weight prediction mechanism: leveraging optimizer-specific update rules to dynamically model weight evolution and accurately predict—prior to forward propagation—the most up-to-date weights for each mini-batch. This is the first approach to strictly guarantee weight consistency and zero staleness per mini-batch under 1F1B, while remaining compatible with arbitrary optimizers. Our method integrates asynchronous scheduling, forward-weight precomputation, and delayed gradient synchronization. Extensive experiments across eight models and three task categories demonstrate throughput comparable to GPipe and PipeDream, while achieving convergence accuracy on par with fully synchronous baselines.

Addresses weight inconsistency in 1F1B schedulesEnsures effective parameter learning in asynchronous trainingPrevents weight staleness across GPUs in pipelines

To address the substantial pipeline bubble overhead, imbalanced GPU memory utilization, and throughput limitations in large-scale neural network training, this paper proposes a synchronous bidirectional pipeline parallelism mechanism. Our approach introduces a novel synchronous bidirectional micro-batch scheduling strategy that achieves dynamic activation memory balancing and minimizes pipeline bubbles while preserving full-precision computation. Integrated with Transformer-specific distributed training optimizations, the proposed method trains a 1.3-billion-parameter GPT-2 model on the Piz Daint supercomputer (2,048 GPUs). It achieves 1.16×–2.34× higher throughput compared to state-of-the-art synchronous and asynchronous pipeline parallel methods. This advancement significantly improves training efficiency and hardware resource utilization for large-scale models.

Efficiently training large-scale neural networks with bidirectional pipelinesImproving training throughput for billion-parameter modelsReducing pipeline bubbles and balancing memory consumption

Latest Papers

What's happening recently
View more

Existing large-scale model training systems struggle to flexibly compose diverse parallelization strategies, often relying on manual expert tuning and lacking generality. This work proposes a programmable distributed training system that enables users to declaratively specify composite parallelism strategies—such as data, pipeline, and expert parallelism—through model annotations and scheduling directives. These specifications are compiled via a unified intermediate representation (IR) into device-level execution plans, fully decoupling strategy definition from runtime execution over a global compute-communication DAG. The system is the first to support automatic compilation of user-defined composite strategies, matching the performance of established approaches like ZeRO while significantly improving both performance and memory efficiency in complex scenarios such as DeepSeek-V3’s DualPipe.

distributed trainingflexibilitymodel parallelism

This work addresses the challenges of training large language models across geographically distributed GPU clusters, where heterogeneous network bandwidth and regional electricity price disparities hinder existing pipeline parallelism strategies from simultaneously minimizing job completion time and power cost, while also suffering from head-of-line blocking. To overcome these limitations, the authors propose BACE-Pipe, a novel framework that jointly models real-time network utilization and job characteristics for the first time. BACE-Pipe integrates dynamic priority scheduling, bandwidth-aware path planning, and electricity-price-aware GPU allocation to co-optimize pipeline execution across regions. Experimental results demonstrate that the approach effectively mitigates head-of-line blocking and significantly reduces both latency and operational cost in multi-tenant environments, achieving 27.9%–64.7% lower average job completion time and 12.6%–30.6% reduction in total electricity expenditure.

Bandwidth HeterogeneityElectricity CostGeo-Distributed Training

This work addresses the degraded convergence in asynchronous pipeline-parallel training caused by parameter inconsistency between forward and backward passes. To mitigate this issue, the paper proposes Asynchronous Multi-directional Pipeline Parallelism (AMDP), which uniquely integrates a pipeline-depth-aware multi-directional concurrency mechanism with bounded parameter mismatch control. AMDP stabilizes training convergence while maintaining high hardware utilization by limiting the number of micro-batches in the initial stage, dynamically scheduling multiple concurrent pipelines, and accumulating gradients across batches. Experimental results demonstrate that AMDP significantly accelerates training for GPT- and BERT-style models while achieving convergence performance comparable to synchronous methods.

asynchronous trainingconvergence degradationlarge-scale models

Existing evaluations of pipeline parallelism scheduling strategies for large language models are limited by analytical models that neglect communication overhead and costly end-to-end experiments. This work proposes a unified evaluation framework that integrates formal modeling, tabular scheduling abstractions, and communication-aware execution simulation, enabling—for the first time—joint modeling of structured schedule representations and communication costs. Using this framework, we systematically compare GPipe, 1F1B, Chimera, and Hanayo across diverse hardware configurations, revealing that scheduling efficacy is highly dependent on the execution environment and challenging the conventional paradigm of relying solely on structural metrics such as bubble ratio. Our experiments show that GPipe and 1F1B yield similar training times, though 1F1B uses less activation memory; Chimera is advantageous only with few microbatches or highly efficient communication; and Hanayo performs well within its applicable scenarios but is sensitive to network bottlenecks.

communication-awareLLM trainingpipeline parallelism

This work addresses the communication latency and network sensitivity caused by frequent gradient synchronization in geographically distributed data-parallel training. The authors propose CPDP, a novel strategy that explicitly models synchronization frequency as a tunable system parameter and introduces a hybrid coordination mechanism combining gradient AllReduce with SlowMo-style parameter averaging. This approach significantly enhances communication efficiency while preserving convergence guarantees. Implemented within a PyTorch DDP-compatible framework, CPDP achieves a 2.44 percentage point improvement in test accuracy over standard DDP on ResNet-50 trained on CIFAR-100 (with K=4), reduces average training time by 13.8%, and cuts synchronization exposure time by approximately 50%. Consistent gains are also demonstrated on ViT-S trained on TinyImageNet.

communication latencydata-parallel trainingdistributed machine learning

Hot Scholars

SY

Si-Yang Liu

Nanjing University
Machine LearningTabular DataLLMs
HJ

Han-Jia Ye

Nanjing University
Machine LearningData MiningMetric LearningMeta-Learning
LY

Lijun Yu

Google DeepMind
Video GenerationMultimodal Foundation Model
MC

Matthieu Cord

Professor Sorbonne University / Scientific Director valeo.ai
Computer VisionImage ProcessingMachine LearningArtificial Intelligence
ML

Mingjie Liu

Assistant Professor, Department of Chemistry, University of Florida
computational materials scienceenergy conversion and storagemachine learningdata science