🤖 AI Summary
This work addresses the degraded convergence in asynchronous pipeline-parallel training caused by parameter inconsistency between forward and backward passes. To mitigate this issue, the paper proposes Asynchronous Multi-directional Pipeline Parallelism (AMDP), which uniquely integrates a pipeline-depth-aware multi-directional concurrency mechanism with bounded parameter mismatch control. AMDP stabilizes training convergence while maintaining high hardware utilization by limiting the number of micro-batches in the initial stage, dynamically scheduling multiple concurrent pipelines, and accumulating gradients across batches. Experimental results demonstrate that AMDP significantly accelerates training for GPT- and BERT-style models while achieving convergence performance comparable to synchronous methods.
📝 Abstract
Pipeline parallelism is essential for large-scale model training, but existing asynchronous approaches often degrade convergence due to parameter mismatch between forward and backward passes. We propose Asynchronous Multi-Directional Pipeline parallelism (AMDP) to mitigate this issue while sustaining high utilization. AMDP limits the first stage of each pipeline to process at most two minibatches before backpropagation, bounding the number of parameter updates between forward and backward passes. To alleviate the resulting pipeline bubbles, AMDP launches multiple concurrent pipelines and adapts their number according to pipeline depth. In addition, AMDP accumulates gradients across minibatches and applies them in a single update, ensuring that only a bounded number of minibatches experience parameter mismatch, limited to within one optimization step. Experiments on GPT- and BERT-style models demonstrate that AMDP significantly accelerates training while preserving convergence.