AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training

📅 2026-05-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the degraded convergence in asynchronous pipeline-parallel training caused by parameter inconsistency between forward and backward passes. To mitigate this issue, the paper proposes Asynchronous Multi-directional Pipeline Parallelism (AMDP), which uniquely integrates a pipeline-depth-aware multi-directional concurrency mechanism with bounded parameter mismatch control. AMDP stabilizes training convergence while maintaining high hardware utilization by limiting the number of micro-batches in the initial stage, dynamically scheduling multiple concurrent pipelines, and accumulating gradients across batches. Experimental results demonstrate that AMDP significantly accelerates training for GPT- and BERT-style models while achieving convergence performance comparable to synchronous methods.
📝 Abstract
Pipeline parallelism is essential for large-scale model training, but existing asynchronous approaches often degrade convergence due to parameter mismatch between forward and backward passes. We propose Asynchronous Multi-Directional Pipeline parallelism (AMDP) to mitigate this issue while sustaining high utilization. AMDP limits the first stage of each pipeline to process at most two minibatches before backpropagation, bounding the number of parameter updates between forward and backward passes. To alleviate the resulting pipeline bubbles, AMDP launches multiple concurrent pipelines and adapts their number according to pipeline depth. In addition, AMDP accumulates gradients across minibatches and applies them in a single update, ensuring that only a bounded number of minibatches experience parameter mismatch, limited to within one optimization step. Experiments on GPT- and BERT-style models demonstrate that AMDP significantly accelerates training while preserving convergence.
Problem

Research questions and friction points this paper is trying to address.

pipeline parallelism
asynchronous training
parameter mismatch
large-scale models
convergence degradation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Asynchronous Pipeline Parallelism
Parameter Mismatch Mitigation
Multi-Directional Pipelining
Gradient Accumulation
Large-Scale Model Training
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Ling Chen
State Key Laboratory of Blockchain and Data Security, Zhejiang University, Hangzhou, China; College of Computer Science and Technology, Zhejiang University, Hangzhou, China
H
Houming Wu
State Key Laboratory of Blockchain and Data Security, Zhejiang University, Hangzhou, China; College of Computer Science and Technology, Zhejiang University, Hangzhou, China
W
Wenjie Yu
State Key Laboratory of Blockchain and Data Security, Zhejiang University, Hangzhou, China; College of Computer Science and Technology, Zhejiang University, Hangzhou, China