HyperParallel-FSDP: Topology-Aware Fully Sharded Training with Layout-Driven Muon on Ascend SuperPods

📅 2026-09-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决分布式训练中的性能问题,提出HyperParallel-FSDP方法,通过双模式DTensor执行、拓扑感知FSDP及布局驱动的分布式Muon优化模型并行计算效率。
📝 Abstract
Declarative SPMD programming uses tensor sharding descriptions to drive distributed execution, separating parallelization from model code. However, the evaluated PyTorch DTensor stack dispatches every operator below autograd, incurring repeated dispatch and metadata costs, while lacking an inexpensive end-to-end validation path. Existing FSDP and distributed Muon implementations also mismatch two-tier supernode topologies: FSDP relies on explicit parameter packing and unpacking, and Muon's whole-matrix orthogonalization conflicts with parameter sharding. We observe that distributed tensors need only express sharding semantics at the tensor API boundary above autograd, allowing differentiation and kernels to operate on plain tensors. Based on this insight, we present HyperParallel-FSDP, featuring: (1) dual-mode DTensor execution, using one sharding plan for both a production mode with one-time layout resolution and no steady-state dispatch overhead, and a validation mode with end-to-end metadata propagation, fail-fast checks, and gradient-equivalence testing; (2) topology-aware FSDP, with zero-copy intra-supernode collectives, fused inter-supernode reduction, and a cross-layer backward pipeline that avoids waits on slow links; and (3) layout-driven distributed Muon, with sharding-derived communication groups, deduplicated orthogonalization, and shape-fused Newton-Schulz iterations. On Atlas 900 A3 SuperPoD, HyperParallel-FSDP scales from 16 dies to 384 cards (768 ranks), sustaining 421k tokens/s for a 505B-parameter MoE while FSDP communication uses 2.9% of step time. It reduces mean step time by 29.7% versus PyTorch FSDP2 and 25.5% versus Megatron DDP, with Pearson correlation above 0.999997 over 1,000 steps. Distributed Muon improves profiler step time by 5.4-16.0% over competing systems. Source code is available at https://atomgit.com/mindspore/hyper-parallel.
Problem

Research questions and friction points this paper is trying to address.

DTensor
FSDP
distributed Muon
supernode topologies
tensor sharding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-mode DTensor execution
Topology-aware FSDP
Layout-driven distributed Muon
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mo Sun
Zhejiang University
Yifan Yao
Yifan Yao
Drexel University
Y
Yanwei Liu
Huawei Technologies Co., Ltd
L
Luobin Liu
Huawei Technologies Co., Ltd
Z
Zhenzhang Yang
Huawei Technologies Co., Ltd
K
Kaisheng Wang
Huawei Technologies Co., Ltd
Xiangyu Meng
Xiangyu Meng
中国石油大学(华东)
HPC AI
C
Chen Li
Huawei Technologies Co., Ltd
X
Xizheng Pang
Huawei Technologies Co., Ltd
H
Huilan Li
Huawei Technologies Co., Ltd
X
Xinglei Xu
Huawei Technologies Co., Ltd
Y
Yushi Cui
Huawei Technologies Co., Ltd
X
Xinyao Lin
Zhejiang University
K
Kaiqi Chen
Zhejiang University
Jie Zhang
Jie Zhang
Zhejiang University
SmartNICNetworked SystemsProcessing In Memory
Zeke Wang
Zeke Wang
Zhejiang University
Machine Learning SystemsSmartNICFPGAGPU
T
Teng Su
Huawei Technologies Co., Ltd