Optimal, Non-pipelined Reduce-scatter and Allreduce Algorithms

📅 2024-10-18
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses efficiency bottlenecks in reduce-scatter and allreduce collective communications within processor networks. We propose non-pipelined algorithms that are both round-optimal (requiring exactly ⌈log₂p⌉ rounds) and volume-optimal (minimizing total data movement). Our approach leverages a circulant graph communication topology and a binary-exchange reduction mechanism, enabling deterministic, synchronous protocols under commutative reduction operators. To our knowledge, this is the first construction of a round-optimal and volume-optimal reduce-scatter algorithm; moreover, within the same unified framework, we derive optimal allreduce and scalable alltoall variants—breaking the limitations of conventional tree- or ring-based topologies. The algorithms strictly conform to MPI standard interfaces and integrate directly into MPI_Reduce_scatter_block, MPI_Reduce_scatter, and MPI_Allreduce. Empirical evaluation demonstrates superior scalability and communication efficiency in distributed training and high-performance computing workloads.

Technology Category

Search and Optimization: Distributed SearchMachine Learning: Distributed Machine Learning & Federated LearningMultiagent Systems: Agent Communication

Application Category

Graph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsSystems and Infrastructure for Web, Mobile and WoT: Experiences and lessons learnt from Web-based algorithms and system deploymentsResponsible Web: Human-perceived consequences of algorithmic deployment on the web
📝 Abstract
The reduce-scatter collective operation in which $p$ processors in a network of processors collectively reduce $p$ input vectors into a result vector that is partitioned over the processors is important both in its own right and as building block for other collective operations. We present a surprisingly simple, but non-trivial algorithm for solving this problem optimally in $lceillog_2 p ceil$ communication rounds with each processor sending, receiving and reducing exactly $p-1$ blocks of vector elements. We combine this with a similarly simple, well-known allgather algorithm to get a volume optimal algorithm for the allreduce collective operation where the result vector is replicated on all processors. The communication pattern is a simple, $lceillog_2 p ceil$-regular, circulant graph also used elsewhere. The algorithms assume the binary reduction operator to be commutative and we discuss this assumption. The algorithms can readily be implemented and used for the collective operations MPI_Reduce_scatter_block, MPI_Reduce_scatter and MPI_Allreduce as specified in the MPI standard. We also observe that the reduce-scatter algorithm can be used as a template for round-optimal all-to-all communication and the collective MPI_Alltoall operation.
Problem

Research questions and friction points this paper is trying to address.

Optimal reduce-scatter algorithm
Volume optimal allreduce operation
Round-optimal all-to-all communication template
Innovation

Methods, ideas, or system contributions that make the work stand out.

Optimal reduce-scatter algorithm
Volume optimal allreduce algorithm
Logarithmic communication rounds
🔎 Similar Papers
No similar papers found.
TU Wien
J
Jesper Larsson Träff
TU Wien, Faculty of Informatics, Institute of Computer Engineering, Research Group Parallel Computing 191-4, Treitlstrasse 3, 5th Floor, 1040 Vienna, Austria