NCCL M2N: A Layout- and Topology-Aware Collective for Distributed Tensor Resharding

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the communication bottleneck in weight resharding caused by tensor layout switching during distributed training. We propose layout- and topology-aware collective communication primitives built upon the NCCL framework. By leveraging NVLink and InfiniBand hardware characteristics alongside an aggregated data movement model, our approach balances load through hierarchical routing and overlaps network transmission with local copying to eliminate traffic amplification, thereby enabling efficient data redistribution between source and target process groups. Experimental results demonstrate that the proposed method achieves a 7.9× speedup in single-layer transmission and a 2.09× acceleration in weight synchronization, ultimately reducing the training step time of DeepSeek-V3 by 12.7%.
📝 Abstract
Distributed training and rollout generation often use different tensor layouts, requiring model weights to be resharded across distinct process groups. This M-to-N redistribution is not directly expressed by standard collectives. Flat direct sends duplicate traffic across destination replicas, while gather-then-broadcast concentrates network injection at one root and transfers data that destinations do not need. For DeepSeek-V3 on 256 GPUs, legacy all-gather plus broadcast accounts for 29.4% of the reported reinforcement-learning step time. We present NCCL M2N, a layout- and topology-aware collective primitive for distributed tensor resharding. Given source and destination meshes and placements, it derives the required transfer regions and a global communication schedule. Its hierarchical route balances source contributions across eligible destination ranks, forwards one copy between destination NVLink domains, and completes local replication over NVLink. Network transfer and local replication overlap, avoiding replica-multiplied source egress. An aggregate data-movement model captures the limits of both stages. We evaluate NCCL M2N on up to 256 GB200 GPUs in an NVL72 cluster with NDR InfiniBand. A single FFN-MoE layer transfer achieves up to 7.9x speedup over flat direct sends (9.8 ms versus 77.3 ms). In a separate 256-GPU DeepSeek-V3 NeMo-RL experiment, NCCL M2N reduces reported weight-sync time from 5.78 s to 2.77 s, a 2.09x speedup over legacy all-gather plus broadcast, and reduces step time by 12.7%.
Problem

Research questions and friction points this paper is trying to address.

tensor resharding
collective communication
distributed training
M-to-N redistribution
tensor layout
Innovation

Methods, ideas, or system contributions that make the work stand out.

collective communication
tensor resharding
topology-aware routing
NVLink
distributed training
🔎 Similar Papers
No similar papers found.
K
Kaushik Kandadi
NVIDIA Corporation
Youngeun Kwon
Youngeun Kwon
NVIDIA Corporation
Sreeram Potluri
Sreeram Potluri
NVIDIA Corporation
C
Ching-Hsiang Chu
NVIDIA Corporation
Ke Wen
Ke Wen
NVIDIA Corporation
P
Pouya Kousha
NVIDIA Corporation
Sangkug Lym
Sangkug Lym
Nvidia
N
Nitin Nitin
NVIDIA Corporation
M
Manjunath Gorentla Venkata
NVIDIA Corporation