Effective MPI: User-defined Datatypes and Cartesian Communicators for Zero-copy All-to-all Communication in Multidimensional Tori

๐Ÿ“… 2026-05-28
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the performance bottleneck of MPI_Alltoall under complex process layouts in high-dimensional torus topologies by proposing a zero-copy all-to-all communication algorithm. The approach decomposes the communication operation dimension by dimension, leveraging MPI Cartesian communicators and user-defined datatypes to implicitly perform data reordering without explicit local memory copies. A double-buffering mechanism and pre-cached dimensional communicators are introduced to enable flexible performance tuning. Experimental results demonstrate that the proposed method matches or even exceeds the performance of native MPI_Alltoall across various high-dimensional torus configurations, revealing significant optimization potential in current MPI implementations.
๐Ÿ“ Abstract
We present and show how to implement a non-trivial all-to-all communication algorithm for arbitrary $d$-dimensional tori effectively in MPI. Given a factorization of the number of processes $p$ into $d$ factors that can be mapped onto a $d$-dimensional torus, we first utilize a Cartesian communicator to split a given $p$-process MPI communicator into, for each MPI process, $d$ smaller communicators spanning each of the dimensions of the torus to which the process belongs, and cache these communicators in order to avoid expensive splitting at each all-to-all operation. The all-to-all operation itself is decomposed into a sequence of $d$ MPI_Alltoall operations on the dimension-wise communicators. The non-trivial data rearrangement before and after each MPI_Alltoall call is implicit only and effected by MPI derived datatypes. This makes the implementation of the algorithm formally \emph{zero-copy}, meaning that no explicit process-local reordering of data blocks ever has to be performed. In order to achieve this, the algorithm employs a double-buffering scheme with modest temporary buffer requirements. By choosing the factorization of $p$ and selecting appropriate implementations for the component MPI_Alltoall operations, the presented implementation gives ample opportunities for algorithm tuning and adaptation to the particular high-performance system. A few, select experimental results show competitive performance with native MPI_Alltoall implementations and illustrate problems that common MPI_Alltoall implementations may have.
Problem

Research questions and friction points this paper is trying to address.

all-to-all communication
multidimensional tori
zero-copy
MPI datatypes
Cartesian communicators
Innovation

Methods, ideas, or system contributions that make the work stand out.

zero-copy
user-defined datatypes
Cartesian communicators
all-to-all communication
multidimensional tori
๐Ÿ”Ž Similar Papers
No similar papers found.