optimize collective communication

Designs, implements, and evaluates algorithms, libraries, and low-level primitives for collective communication (e.g., broadcast, reduce, allreduce, gather, scatter) in parallel and distributed systems; analyzes and traces collective operations and tunes algorithmic choices and implementation parameters to optimize latency, bandwidth, scalability, and resource usage across different network topologies and system hierarchies.

optimizecollectivecommunication

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.03
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$216K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Collective Communication Profiling of Modern-day Machine Learning Workloads

Jul 03, 2025
JG
Jit Gupta
🏛️ Juniper Networks | Stanford University

In large-scale distributed ML training, collective communication primitives (e.g., AllReduce, AllGather, Broadcast) generate high-bandwidth, bursty traffic, causing network congestion and packet loss. Method: This work systematically characterizes the communication behavior of mainstream LLMs—including DeepSeek-V3, GPT, and Llama—under diverse parallelism strategies, scales, and network topologies. Leveraging fine-grained empirical analysis of NVIDIA NCCL logs, we quantitatively identify how operation type, message size, and request distribution affect network anomalies. We propose a resource-coordinated optimization framework tailored to LLM communication patterns, jointly optimizing collective primitive scheduling and network topology adaptation. Contribution/Results: Experiments demonstrate that our approach significantly mitigates congestion, improving both communication efficiency and stability in distributed training and inference—without modifying model architecture or training algorithms.

Analyze bursty traffic patterns in ML collective communicationImprove collective communication frameworks for network anomaliesOptimize network resources for diverse ML workloads

Optimal, Non-pipelined Reduce-scatter and Allreduce Algorithms

Oct 18, 2024
JL
Jesper Larsson Träff
🏛️ TU Wien

This work addresses efficiency bottlenecks in reduce-scatter and allreduce collective communications within processor networks. We propose non-pipelined algorithms that are both round-optimal (requiring exactly ⌈log₂p⌉ rounds) and volume-optimal (minimizing total data movement). Our approach leverages a circulant graph communication topology and a binary-exchange reduction mechanism, enabling deterministic, synchronous protocols under commutative reduction operators. To our knowledge, this is the first construction of a round-optimal and volume-optimal reduce-scatter algorithm; moreover, within the same unified framework, we derive optimal allreduce and scalable alltoall variants—breaking the limitations of conventional tree- or ring-based topologies. The algorithms strictly conform to MPI standard interfaces and integrate directly into MPI_Reduce_scatter_block, MPI_Reduce_scatter, and MPI_Allreduce. Empirical evaluation demonstrates superior scalability and communication efficiency in distributed training and high-performance computing workloads.

Optimal reduce-scatter algorithmRound-optimal all-to-all communication templateVolume optimal allreduce operation

This study addresses collective communication bottlenecks in distributed large language models operating across heterogeneous hardware and hybrid parallelism. We propose a three-tier optimization framework centered on collective communication that integrates a communication taxonomy with topology-aware scheduling, dynamic GPU runtime mapping, and computation-communication overlap. This approach achieves system-level optimization spanning planning, execution, and end-to-end coordination. Furthermore, this work establishes a generalized optimization paradigm tailored for multi-NIC heterogeneous interconnects, significantly enhancing both training and inference efficiency. By systematically addressing these challenges, the proposed framework not only improves performance in complex distributed environments but also delineates critical future research directions for communication optimization in large-scale distributed systems.

Collective CommunicationCommunication PlanningComputation Coordination

Efficient Direct-Connect Topologies for Collective Communications

Feb 07, 2022
LZ
Liangyu Zhao
🏛️ University of Washington | Raytheon BBN | MIT

Existing direct-connect network topologies exhibit inadequate adaptability to diverse scales, node degrees, and latency–bandwidth trade-offs in high-performance computing (HPC) collective communication. Method: This paper proposes an iterative expansion framework grounded in small-scale optimal base topologies. It introduces, for the first time, a graph-synthesis-driven automatic topology generation mechanism and designs the first polynomial-time collective communication scheduling algorithm for canonical large-scale topologies—including Dragonfly and Fat-Tree. Contribution/Results: The work unifies topology synthesis and scheduling optimization within a single modeling framework, enabling cross-platform deployment and large-scale simulation validation. Experimental evaluation demonstrates that the proposed approach reduces average communication latency by 23.6% and improves bandwidth utilization by 31.4% compared to conventional topologies, while significantly enhancing scalability and practical applicability in real-world HPC systems.

Information FlowNetwork StructureOptimization

Demystifying NCCL: An In-depth Analysis of GPU Communication Protocols and Algorithms

Jul 07, 2025
ZH
Zhiyi Hu
🏛️ ETH Zürich | NVIDIA Corporation | Broadcom Inc.

NCCL—the de facto standard library for high-performance collective communication in GPU clusters—lacks transparency in its internal protocol selection, channel orchestration, and cross-node memory movement mechanisms, hindering systematic performance analysis and bottleneck identification. Method: We present the first systematic reverse-engineering of NCCL’s multi-tier communication architecture, integrating trace-driven modeling with fine-grained analysis of its three core protocols (Simple, LL, LL128) to uncover the dynamic scheduling logic of ring and tree algorithms—and their associated data movement strategies—under realistic AI training workloads. Contribution/Results: Based on these insights, we develop ATLAHS, a reproducible, industrial-grade simulation toolchain that accurately models NCCL’s communication behavior, enabling precise performance prediction and root-cause bottleneck diagnosis. Our work establishes a verifiable theoretical foundation and practical toolset for designing high-performance communication libraries, optimizing AI training systems, and enabling hardware-software co-tuning.

Analyze NCCL's opaque internal design and protocolsOptimize collective communication in large-scale AI trainingUnderstand GPU communication channels and memory handling

Latest Papers

What's happening recently
View more

This work addresses the performance bottleneck in large-scale AI training and inference caused by inefficient overlap between computation and communication, as well as high communication overhead. To this end, the authors develop a customized collective communication library for the Meta MTIA 300 accelerator, integrating a backend network within the chip package for the first time. By combining near-memory computing (NMC) with a dedicated message engine (ME), the design enables full communication offload. The paper introduces a compiler-driven communication model, topology-aware algorithms, and one-sided communication primitives tailored for inference, optimizing collective operations across heterogeneous scale-up and scale-out networks. Experiments demonstrate that, in training scenarios, intra-rack collective bandwidth reaches 940 GB/s with less than 0.5% impact on concurrent compute throughput; in inference, communication latency is significantly reduced, greatly enhancing compute-communication pipeline efficiency.

acceleratorcollective communicationcompute-communication overlap

Collective communication in distributed machine learning often becomes a performance bottleneck due to the neglect of physical network topology and process group structure. This work proposes a scalable and general framework for synthesizing collective communication algorithms that, for the first time, incorporates process-group awareness into algorithm generation, supporting arbitrary communication patterns. By integrating topology-aware modeling with optimized search strategies, the framework automatically generates high-performance communication algorithms tailored to the actual process groups and underlying network topology. Experimental results demonstrate that the framework can synthesize an All-to-All algorithm for a 512-NPU system within 11.68 minutes, achieving performance close to the theoretical optimum.

algorithm synthesiscollective communicationdistributed machine learning

This work addresses the high complexity of existing high-performance computing (HPC) performance analysis tools, which hinders students’ intuitive understanding of parallel program performance issues. To bridge this gap, the paper introduces EduMPI—the first educational tool that integrates HPC cluster operations and MPI performance analysis within a streamlined graphical interface. EduMPI enables near real-time, physically node-layout-aware communication visualization, facilitating interactive identification of load imbalance and other performance bottlenecks. User studies demonstrate that, compared to professional-grade tools, EduMPI significantly lowers the learning barrier and effectively enhances students’ comprehension of parallel performance characteristics, thereby improving the practicality and accessibility of parallel programming education.

educational integrationMPIparallel programming education

This work addresses the challenge of simultaneously minimizing worst-case communication overhead and computational load for general functions admitting d-ary decompositions in distributed computing. The authors propose a deterministic Interweaved Clique (IC) assignment framework grounded in combinatorial design theory. This approach circumvents the restrictive existence conditions of Steiner systems, thereby revealing for the first time the fundamental scaling laws of the problem over a significantly broader range of parameters, while permitting modest heterogeneity in workers’ storage loads. The constructed IC scheme achieves communication cost within a constant factor of 4e from the information-theoretic lower bound and maintains order-wise optimal computation load.

communication costcomputation loaddistributed computing

This work addresses the limitations of existing performance evaluation approaches for distributed computing continua, which often focus on a single dimension and fail to holistically characterize the behavior of cross-layer heterogeneous systems. The paper presents the first systematic framework that establishes a comprehensive taxonomy of performance metrics spanning three layers—computation, networking, and application/user—as well as emerging non-functional attributes such as sustainability and observability. By integrating mathematical modeling with cross-layer analysis, the study rigorously defines the applicability, measurement phases, and specifications for each metric category. The resulting framework is both clearly structured and extensible, offering a solid theoretical foundation and practical guidance for unified performance assessment in dynamic, heterogeneous environments.

Cross-layer MetricsDistributed Computing ContinuumHeterogeneous Systems

Hot Scholars

HL

Haibin Lin

Bytedance
Machine Learning SystemsNatural Language Processing
SS

Siddharth Singh

Research Scientist at Nvidia
High Performance ComputingArtificial Intelligence
EP

Evangelos Pournaras

Professor of Trustworthy Distributed Intelligence, UKRI Future Leaders Fellow, University of Leeds
distributed intelligenceblockchaincollective decision-makingdigital democracy
XC

Xiaowen Chu

IEEE Fellow, Professor, Data Science and Analytics, HKUST(GZ)
GPU ComputingMachine Learning SystemsParallel and Distributed ComputingWireless Networks
QA

Quentin Anthony

PhD Student, Ohio State University
HPCDeep LearningParallel Computing