Score
Designs and implements schedulers, protocols, and runtime mechanisms that plan, order, and tune message exchanges—including collective and MPI-style operations—by coalescing or fragmenting messages, overlapping communication with computation, and applying low-bit aggregation or reducer-capacity checks to reduce latency, bandwidth use, and synchronization overhead. Builds measurement and profile-guided tooling and analyzes message-passing patterns, routing and placement decisions, and communication-aware partitioning to drive runtime scheduling and protocol configuration.
This work addresses the complexity, routing ambiguity, and unreliable shutdown commonly introduced by ad hoc glue code in existing modular distributed systems. To overcome these issues, the paper proposes CNS—a lightweight, local-first hybrid event bus that seamlessly bridges local and distributed publish-subscribe contexts through a unified event model and consistent routing semantics. CNS employs an asynchronous fire-and-forget primary path while supporting request-response extensions on the same topic. It integrates typed event keys, family-based serialization and validation, and NATS-backed distributed transport. Prototype evaluation demonstrates low-latency performance: approximately 30 microseconds for local delivery, 1.26–1.37 milliseconds for purely distributed communication, and 1.64–1.89 milliseconds for hybrid bridging, with validation overhead remaining manageable—making CNS suitable for structured inter-process communication and efficient messaging among resource-constrained nodes.
Traditional MPI C interfaces lack type safety and generic programming support, hindering the adoption of modern C++ in high-performance computing (HPC). To address this, we propose a layout-agnostic message-passing abstraction that—leveraging the Noarr library—introduces first-class data layout and traversal abstractions into the MPI communication layer for the first time. Our approach decouples communication semantics from memory layout while enabling type-safe, generic, and composable MPI interfaces. By tightly integrating modern C++ template metaprogramming with low-level MPI mechanisms, our design preserves near-native MPI performance while significantly improving interface safety, expressiveness, and modularity. We evaluate the framework using distributed GEMM as a case study: results demonstrate enhanced code reusability, greater development flexibility, and zero computational overhead relative to baseline MPI implementations.
To address the challenge of fine-grained characterization and cross-application comparison of MPI communication behavior in HPC applications, this paper introduces, for the first time within the Caliper performance profiling framework, a “Communication Region” mechanism. This mechanism annotates MPI call boundaries and associates them with process- and data-level metrics to enable context-aware quantification of communication overhead. Leveraging the Benchpark benchmark suite and the Thicket analysis library, we model and visualize canonical communication patterns—including halo exchanges—across AMG2023, Kripke, and Laghos on both CPU and GPU platforms. Our approach supports quantitative cross-application message volume analysis, scalability divergence attribution, and precise bottleneck identification. It significantly improves the accuracy and comparability of MPI communication behavior analysis. The method has been validated on real-world simulation codes, demonstrating both effectiveness and practical utility.
In Time-Sensitive Networking (TSN), multicast communication improves bandwidth efficiency but exacerbates scheduling complexity due to port contention and queue resource constraints. Method: This paper proposes a fine-grained multicast tree partitioning approach that dynamically decomposes large multicast trees into smaller multicast or unicast subtrees, integrating time-triggered planning, adaptive tree-splitting algorithms, and multi-strategy scheduling to jointly optimize resource utilization and schedulability under heterogeneous topologies. Contribution/Results: To the best of our knowledge, this is the first systematic incorporation of multicast partitioning into TSN time-triggered flow scheduling. The method guarantees end-to-end latency bounds and achieves load balancing across switches. Experimental evaluation shows a 5–15% reduction in flow rejection rate and a 5–125% increase in throughput compared to an unpartitioned baseline, significantly improving flow admission ratio and scheduling feasibility in dynamic TSN environments.
This work addresses the performance bottleneck of MPI_Alltoall under complex process layouts in high-dimensional torus topologies by proposing a zero-copy all-to-all communication algorithm. The approach decomposes the communication operation dimension by dimension, leveraging MPI Cartesian communicators and user-defined datatypes to implicitly perform data reordering without explicit local memory copies. A double-buffering mechanism and pre-cached dimensional communicators are introduced to enable flexible performance tuning. Experimental results demonstrate that the proposed method matches or even exceeds the performance of native MPI_Alltoall across various high-dimensional torus configurations, revealing significant optimization potential in current MPI implementations.
This work addresses the scalability bottleneck in MPI initialization caused by the reliance on the global communicator MPI_COMM_WORLD in traditional MPI implementations, particularly at exascale. By rearchitecting MPICH’s internal design, the authors present the first mainstream MPI implementation that fully decouples MPI_COMM_WORLD and introduces a compliant, true Sessions model aligned with the MPI-4 standard. The proposed approach employs explicit hierarchical process-set management and a scalable initialization protocol, substantially improving startup scalability. Experimental results demonstrate that the new mechanism efficiently supports exascale-class supercomputing systems, establishing a critical foundation for deploying MPI on next-generation ultra-large-scale platforms.
This study investigates the scalability and performance of process and thread schedulers under memory-intensive workloads in multi-core shared-memory systems, focusing on a 3D tensor row-sorting task. The authors design and evaluate several scheduling strategies: on the thread side, an AIMD-based adaptive chunking mechanism inspired by TCP congestion control is introduced, coupled with exponential weighted moving average to dynamically adjust concurrency; on the process side, a bounded prolific/collective model is employed alongside one-to-one, one-to-many, and many-to-many pipelined communication patterns to enable flexible task distribution. Experimental results on a 24-core x86-64 platform demonstrate that thread-level scheduling consistently outperforms process-level scheduling, with dynamic and guided strategies achieving the best performance, while the many-to-many pipeline exhibits superior scalability for large-scale tasks.
This work addresses the inefficiency of MPI_Alltoallv in irregular communication scenarios, where repeated metadata processing degrades performance. For the first time, we introduce a persistent Remote Memory Access (RMA) mechanism into Alltoallv, decoupling initialization from execution to enable reuse of communication metadata and window state. We systematically evaluate the performance trade-offs between fence- and lock-based synchronization strategies. The proposed approach supports hierarchical scalability and demonstrates clear advantages at 448 processes with message sizes of 32 KB or larger. In large-message regimes, execution time improves from 2.49 seconds to 1.54 seconds, achieving up to a 44% speedup. This significantly enhances the scalability and practicality of irregular all-to-all communication in large-scale high-performance computing environments.
This work addresses the challenges of workflow task composition in high-throughput, petabyte-scale data processing environments, where resource heterogeneity and execution overhead significantly impact performance. The authors propose a hybrid task composition strategy that dynamically balances task independence against execution grouping, formulated within a multi-objective optimization framework to achieve Pareto-optimal trade-offs among throughput, I/O cost, and CPU efficiency. Leveraging workflow DAG modeling and high-dimensional parameter space simulation, the approach enables policy-driven automated synthesis of workflows. Experimental results demonstrate that the proposed strategy achieves up to a 3.8× improvement in throughput and reduces network overhead by as much as 14.9× compared to baseline methods, offering a scalable workflow synthesis framework for extreme-scale scientific computing.