StrataCL: Fabric-Native Communication Library for Production Supernodes

๐Ÿ“… 2026-07-28
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses communication bottlenecks in distributed AI training, which often stem from the separation between user buffers and communication buffers, leading to redundant data copies or high registration overhead. To overcome this, the authors propose a native Fabric communication library tailored for supernode architectures, featuring an innovative โ€œallocate-and-registerโ€ mechanism that enables direct communication from user buffers without data copying. The design further integrates NPU-core load-balanced partitioning with NPU-driven SDMA offloading to co-optimize communication operators holistically. Evaluated on Huawei CloudMatrix384, the approach achieves a 1.6ร— improvement in collective communication bandwidth and a 1.4ร— increase in MoE scheduling/merging bandwidth. It also delivers a 1.9ร— higher LLM inference throughput, a 2.2ร— reduction in P99 TTFT latency, and 1.4ร— and 1.3ร— faster training iteration times for LLMs and recommendation systems, respectively.
๐Ÿ“ Abstract
Modern distributed AI workloads run across hundreds of accelerators, making communication a major bottleneck. Existing communication libraries remain largely buffer-centric because user and communication buffers are managed separately, causing redundant data copies or costly user-buffer registration. This paper presents StrataCL, a zero-redundancy and fabric-native communication library for production supernodes. StrataCL introduces registration-on-allocation to realize user-buffer direct communication, and designs communication operators with workload-balanced NPU-core partitioning and NPU-driven SDMA offloading to exploit supernode architecture features. On the Huawei CloudMatrix384, StrataCL improves collective bus bandwidth by up to 1.6x and improves MoE dispatch/combine bus bandwidth by up to 1.4x. Across three production workloads, StrataCL improves LLM inference throughput by 1.9x, reduces P99 TTFT by 2.2x, and reduces LLM and Recsys training iteration time by 1.4x and 1.3x, respectively.
Problem

Research questions and friction points this paper is trying to address.

communication bottleneck
buffer management
distributed AI
redundant data copies
user-buffer registration
Innovation

Methods, ideas, or system contributions that make the work stand out.

StrataCL
fabric-native
registration-on-allocation
NPU-driven SDMA offloading
zero-redundancy communication
๐Ÿ”Ž Similar Papers
No similar papers found.
Tiancheng Hu
Tiancheng Hu
University of Cambridge
natural language processingcomputational social science
J
Jin Qin
SKLP, Institute of Computing Technology, CAS; University of Chinese Academy of Sciences
Yuzheng Wang
Yuzheng Wang
Fudan University
Knowledge DistillationVision-Language ModelAIGC
Ke Liu
Ke Liu
Institute of Computing Technology, Chinese Academy of Sciences
TCP/IPSoftware-defined Network(SDN)Datacenter NetworkOperating SystemDistributed System
T
TangShengsheng Li
Shanghai Jiao Tong University
S
Sheng Wang
Shanghai Jiao Tong University
Z
Zhongzhe Hu
Shanghai Jiao Tong University
Tianlun Hu
Tianlun Hu
Huawei Technologies Co., Ltd.
LLMReinforcement LearningTransfer LearningNetwork SlicingImage Processing
Wei Wang
Wei Wang
Shanghai Jiao Tong University
speech recognitionspeech enhancementtext-to-speech
L
Lijun Li
Shanghai Jiao Tong University
J
Jingbin Zhou
Shanghai Jiao Tong University
X
Xiaoming Bao
Shanghai Jiao Tong University
H
Hongwei Sun
Shanghai Jiao Tong University
Jieru Zhao
Jieru Zhao
Associate Professor, Shanghai Jiao Tong University
Hardware-software co-designAI acceleration and systemCompilerFPGAHigh-level synthesis
H
Huimin Cui
SKLP, Institute of Computing Technology, CAS; University of Chinese Academy of Sciences
Tao Xie
Tao Xie
Peking University Chair Professor, Fudan University Adjunct Top-Talent Professor
Software EngineeringSoftware TestingSoftware AnalyticsMining Software Repositories
Chenxi Wang
Chenxi Wang
Institute of Computing Technology, Chinese Academy of Sciences
Operating SystemManaged RuntimeProgramming Language