HCCL: Collective Communication for Meta Training and Inference Accelerators

📅 2026-07-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the performance bottleneck in large-scale AI training and inference caused by inefficient overlap between computation and communication, as well as high communication overhead. To this end, the authors develop a customized collective communication library for the Meta MTIA 300 accelerator, integrating a backend network within the chip package for the first time. By combining near-memory computing (NMC) with a dedicated message engine (ME), the design enables full communication offload. The paper introduces a compiler-driven communication model, topology-aware algorithms, and one-sided communication primitives tailored for inference, optimizing collective operations across heterogeneous scale-up and scale-out networks. Experiments demonstrate that, in training scenarios, intra-rack collective bandwidth reaches 940 GB/s with less than 0.5% impact on concurrent compute throughput; in inference, communication latency is significantly reduced, greatly enhancing compute-communication pipeline efficiency.
📝 Abstract
We present HCCL, a collective communication library co-designed with Meta's MTIA 300 accelerator, the first Meta chip to integrate backend networking directly on chip package. MTIA 300 includes dedicated message engines (MEs) with near-memory compute (NMC) that fully offload collective execution from the compute grid, enabling large overlap between computation and communication. HCCL uses a compiled communication model in which the host generates a complete description of each collective including dependencies. We describe the control and data path architecture, topology-aware algorithm selection across MTIA 300's asymmetric scale-up and scale-out network, and optimizations for both training and inference workloads. For training, HCCL achieves up to 940 GB/s on intra-rack collectives while introducing less than 0.5% degradation to concurrent compute throughput. For inference, we leverage one-sided communication primitives that bypass the scheduling path to minimize collective latency and describe collective designs that improve compute-communication pipelining for latency-sensitive workloads.
Problem

Research questions and friction points this paper is trying to address.

collective communication
accelerator
training
inference
compute-communication overlap
Innovation

Methods, ideas, or system contributions that make the work stand out.

collective communication
near-memory compute
compiled communication model
topology-aware algorithm
one-sided communication
🔎 Similar Papers
2024-06-07International Symposium on High-Performance Computer ArchitectureCitations: 5
Wesley Bland
Wesley Bland
Meta
Fault ToleranceMessage Passing InterfaceDistributed ComputingParallel Computing
T
Tiago Antunes
Meta Platforms
L
Lars Paul Huse
Meta Platforms
C
Chidambaram Muthu
Meta Platforms
A
Adel Abouchaev
Meta Platforms
R
Rabib Alam
Meta Platforms
A
Abdullah Alperen
Meta Platforms
A
Alexey Andronov
Meta Platforms
J
Jose Anto Akkara
Meta Platforms
V
Vineet Badhwar
Meta Platforms
Pavan Balaji
Pavan Balaji
Argonne National Laboratory
Parallel and Distributed Computing
D
Daniel Berkovitch
Meta Platforms
B
Bartosz Bogdanski
Meta Platforms
S
Shmeelok Chakraborty
Meta Platforms
Sungjun Cho
Sungjun Cho
PhD Student, University of Wisconsin-Madison
Machine LearningNatural Language ProcessingGeometric Deep LearningProbabilistic Modeling
John Choi
John Choi
Postdoctoral Researcher, New York University
neural engineering
J
James Custer
Meta Platforms
R
Rodrigo De Castro
Meta Platforms
Nguyen Dinh Pham
Nguyen Dinh Pham
University of Houston
M
Matthew Edwards
Meta Platforms
Kristian Evensen
Kristian Evensen
Simula Research Laboratory
NetworksMultihomingBandwidth aggregationMultilinkVideo Streaming
E
Evan Ezell
Meta Platforms
A
Alex Finestead
Meta Platforms
Seth Goldstein
Seth Goldstein
Northwestern University Lurie Children's Hospital
Pediatric Surgery
P
Prankur Gupta
Meta Platforms