Analyzing Persistent Alltoallv RMA Implementations for High-Performance MPI Communication

📅 2026-04-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the inefficiency of MPI_Alltoallv in irregular communication scenarios, where repeated metadata processing degrades performance. For the first time, we introduce a persistent Remote Memory Access (RMA) mechanism into Alltoallv, decoupling initialization from execution to enable reuse of communication metadata and window state. We systematically evaluate the performance trade-offs between fence- and lock-based synchronization strategies. The proposed approach supports hierarchical scalability and demonstrates clear advantages at 448 processes with message sizes of 32 KB or larger. In large-message regimes, execution time improves from 2.49 seconds to 1.54 seconds, achieving up to a 44% speedup. This significantly enhances the scalability and practicality of irregular all-to-all communication in large-scale high-performance computing environments.

Technology Category

Data Mining & Knowledge Management: Scalability, Parallel & Distributed SystemsMachine Learning: Scalability of ML SystemsMultiagent Systems: Agent Communication

Application Category

Security and Privacy: Large-scale security measurementsSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesSystems and Infrastructure for Web, Mobile and WoT: Virtualization and resource management in Web systems and infrastructures
📝 Abstract
Collective communication operations such as MPI_Alltoallv are central to many HPC applications, particularly those with irregular message sizes. We design, implement, and evaluate persistent MPI RMA variants of Alltoallv based on fence and lock synchronization, separating a one time initialization phase from per iteration execution to enable reuse of communication metadata and window state across repeated epochs. Our benchmarks tested on LLNL's Dane supercomputer show that the fence-persistent variant consistently outperforms the non-persistent baseline for large message sizes, achieving up to 44% reduction in runtime and improving scalability with increasing process counts; at 448 processes the runtime decreases from 2.49s to 1.54s (38% faster). We further evaluate the algorithms under irregular sparse communication patterns and compare fence- and lock-based designs, including hierarchical extensions. Message-size sweeps and a break-even model demonstrate that persistence provides immediate payoff for messages greater or equal to 32,768 bytes, while smaller messages show limited benefit due to metadata amortization costs. These results indicate that persistent RMA Alltoallv is a practical approach for workloads with large messages, where removing repeated metadata processing leaves runtime dominated by data movement, as evidenced by the increasing time savings with message size, and they clarify the trade-offs between fence and lock synchronization on modern HPC systems.
Problem

Research questions and friction points this paper is trying to address.

MPI_Alltoallv
persistent communication
RMA
collective communication
high-performance computing
Innovation

Methods, ideas, or system contributions that make the work stand out.

persistent MPI
RMA Alltoallv
fence synchronization
lock synchronization
communication metadata reuse
🔎 Similar Papers
No similar papers found.