MegaGraph: Towards Efficient Training of Large-Scale Graph Transformers with Automated Hybrid Parallelism

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the memory bottlenecks and load imbalance caused by attention matrices and embedding layers during large-scale Graph Transformer (GT) training by proposing the first automated hybrid parallelism framework tailored for GTs. The method introduces three specialized strategies: graph-aware context parallelism, heterogeneous pipeline parallelism, and hybrid data parallelism. Leveraging a Profile-Model-Search workflow coupled with an accurate cost model, it efficiently identifies optimal configurations within an exponential search space. Experimental results demonstrate that the proposed framework reduces peak memory consumption by 77.8% and accelerates training by 4.51×. Furthermore, it successfully completes training while preserving accuracy on large-scale graphs where baseline methods fail due to out-of-memory errors.
📝 Abstract
Graph Transformers (GTs) offer superior representation capabilities by overcoming the depth limitations and over-smoothing issues of traditional Graph Neural Networks (GNNs). However, scaling GTs to large graphs poses critical bottlenecks. Specifically, the attention score matrix and its associated topology-aware bias matrix jointly incur significant per-layer memory overhead, and heavy graph embedding layers result in severe workload imbalances. These characteristics are unique to GT training and are not addressed by parallelism techniques designed for either conventional GNNs or Transformers, making a dedicated solution necessary. This paper introduces MegaGraph, the first automated hybrid parallelism framework designed for efficient GT training. MegaGraph designs three specialized strategies, namely graph-aware context parallelism, heterogeneous pipeline parallelism, and hybrid data parallelism, to support efficient training on large-scale graphs. However, coordinating these three parallelism strategies yields an exponentially large configuration space. To address this complexity, an automatic search engine leverages precise cost models via a Profile - Model - Search workflow to identify the optimal parallelism configuration. Evaluations demonstrate that MegaGraph enables training on large-scale graphs where state-of-the-art baselines fail due to out-of-memory (OOM) errors. The framework reduces per-device peak memory by up to 77.8\% and achieves up to 4.51$\times$ training speedup while maintaining model accuracy.
Problem

Research questions and friction points this paper is trying to address.

Graph Transformers
Large-scale graphs
Memory overhead
Workload imbalance
Hybrid parallelism
Innovation

Methods, ideas, or system contributions that make the work stand out.

Graph Transformers
Automated Hybrid Parallelism
Context Parallelism
Pipeline Parallelism
Cost Model
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Tong Qiao
Tong Qiao
Associate Professor, School of Cyberspace, Hangzhou Dianzi University
Media ForensicsAI SecurityDeepFake DetectionData Hiding
A
Ao Zhou
School of Software, Beihang University, China
Y
Yingjie Qi
School of Computer Science and Engineering, Beihang University, China
C
Chunming Hu
School of Software, Beihang University, China; State Key Laboratory of Complex and Critical Software Environment, Beihang University, China
Jianlei Yang
Jianlei Yang
Beihang University
Deep LearningComputer ArchitectureNueromorphic ComputingSpitronicsEDA/VLSI