JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials

πŸ“… Unknown Date
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the efficiency bottleneck in pipeline parallelism for conservative machine learning interatomic potentials (MLIPs), which stems from their dual-backward execution mode’s incompatibility with existing distributed training frameworks. To overcome this challenge, we propose JanusPipe, the first dedicated three-dimensional parallel (pipeline, data, and gradient parallelism) training system tailored for conservative MLIPs. JanusPipe introduces SymFold to enable memory-efficient pipeline parallelism and designs the WaveK dynamic scheduling mechanism to balance computational load across four distinct stages, thereby minimizing pipeline bubbles. Experimental results demonstrate that JanusPipe achieves an average throughput improvement of 1.51Γ— over 1F1B and 1.45Γ— over Hanayo on 32 GPUs, significantly enhancing scaling efficiency.
πŸ“ Abstract
Discovering atom-level phenomena requires molecular dynamics (MD) simulations with ab initio accuracy. Machine learning interatomic potentials (MLIPs) enable stable, high-accuracy MD simulations, and their models exhibit scaling-law trends similar to large language models. However, the lack of scalable and efficient distributed training systems for conservative MLIPs makes them difficult to scale. This is because conservative MLIPs inherently follow a double-backward execution pattern, which involves computing gradients during the forward pass. This pattern creates a mismatch with existing distributed training systems, especially for pipeline parallelism. Therefore, we present JanusPipe, an efficient 3D-parallel (PP/DP/GP) training system tailored for conservative MLIPs. It integrates SymFold to enable memory-efficient pipeline parallelism for conservative MLIPs, and WaveK to reduce pipeline bubbles by balancing the four-phase compute time. Experimental results on 32 GPUs show that JanusPipe improves throughput by $1.51\times$ and $1.45\times$ on average over 1F1B and Hanayo, respectively.
Problem

Research questions and friction points this paper is trying to address.

machine learning interatomic potentials
conservative MLIPs
pipeline parallelism
distributed training
double-backward execution
Innovation

Methods, ideas, or system contributions that make the work stand out.

pipeline parallelism
machine learning interatomic potentials
double-backward execution
SymFold
WaveK
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Hongyu Wang
Hongyu Wang
Institute of Computing Technology, Chinese Academy of Sciences
Deep LearningNatural Language ProcessingComputer Vision
W
Weijian Liu
SKLP, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China; University of Chinese Academy of Sciences, Beijing, China
Hongtao Xu
Hongtao Xu
Fudan Univeristy
Professor
Y
Yan Wang
SKLP, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China; University of Chinese Academy of Sciences, Beijing, China
Mingzhen Li
Mingzhen Li
Institute of Computing Technology, Chinese Academy of Sciences
HPCAI System
W
Weile Jia
SKLP, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
G
Guangming Tan
SKLP, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China; University of Chinese Academy of Sciences, Beijing, China