optimize communication scheduling

Designs and implements schedulers, protocols, and runtime mechanisms that plan, order, and tune message exchanges—including collective and MPI-style operations—by coalescing or fragmenting messages, overlapping communication with computation, and applying low-bit aggregation or reducer-capacity checks to reduce latency, bandwidth use, and synchronization overhead. Builds measurement and profile-guided tooling and analyzes message-passing patterns, routing and placement decisions, and communication-aware partitioning to drive runtime scheduling and protocol configuration.

optimizecommunicationscheduling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.31
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$202K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the complexity, routing ambiguity, and unreliable shutdown commonly introduced by ad hoc glue code in existing modular distributed systems. To overcome these issues, the paper proposes CNS—a lightweight, local-first hybrid event bus that seamlessly bridges local and distributed publish-subscribe contexts through a unified event model and consistent routing semantics. CNS employs an asynchronous fire-and-forget primary path while supporting request-response extensions on the same topic. It integrates typed event keys, family-based serialization and validation, and NATS-backed distributed transport. Prototype evaluation demonstrates low-latency performance: approximately 30 microseconds for local delivery, 1.26–1.37 milliseconds for purely distributed communication, and 1.64–1.89 milliseconds for hybrid bridging, with validation overhead remaining manageable—making CNS suitable for structured inter-process communication and efficient messaging among resource-constrained nodes.

event fabricinter-process communicationmessage routing

Layout-Agnostic MPI Abstraction for Distributed Computing in Modern C++

Oct 19, 2025
JK
Jiří Klepl
🏛️ Charles University

Traditional MPI C interfaces lack type safety and generic programming support, hindering the adoption of modern C++ in high-performance computing (HPC). To address this, we propose a layout-agnostic message-passing abstraction that—leveraging the Noarr library—introduces first-class data layout and traversal abstractions into the MPI communication layer for the first time. Our approach decouples communication semantics from memory layout while enabling type-safe, generic, and composable MPI interfaces. By tightly integrating modern C++ template metaprogramming with low-level MPI mechanisms, our design preserves near-native MPI performance while significantly improving interface safety, expressiveness, and modularity. We evaluate the framework using distributed GEMM as a case study: results demonstrate enhanced code reusability, greater development flexibility, and zero computational overhead relative to baseline MPI implementations.

Enabling flexible distributed computing without performance lossModernizing MPI with C++ features for type safetyProviding layout-agnostic design for distributed applications

Leveraging Caliper and Benchpark to Analyze MPI Communication Patterns: Insights from AMG2023, Kripke, and Laghos

Jul 30, 2025
GN
Grace Nansamba
🏛️ Tennessee Tech University | Lawrence Livermore National Laboratory | University of New Mexico

To address the challenge of fine-grained characterization and cross-application comparison of MPI communication behavior in HPC applications, this paper introduces, for the first time within the Caliper performance profiling framework, a “Communication Region” mechanism. This mechanism annotates MPI call boundaries and associates them with process- and data-level metrics to enable context-aware quantification of communication overhead. Leveraging the Benchpark benchmark suite and the Thicket analysis library, we model and visualize canonical communication patterns—including halo exchanges—across AMG2023, Kripke, and Laghos on both CPU and GPU platforms. Our approach supports quantitative cross-application message volume analysis, scalability divergence attribution, and precise bottleneck identification. It significantly improves the accuracy and comparability of MPI communication behavior analysis. The method has been validated on real-world simulation codes, demonstrating both effectiveness and practical utility.

Analyzing communication patterns in AMG2023, Kripke, and Laghos applicationsEnhancing Caliper to capture MPI communication metrics and statisticsIdentifying communication bottlenecks and scalability differences in HPC systems

Multicast-partitioning in Time-triggered Stream Planning for Time-Sensitive Networks

Oct 23, 2025
HG
Heiko Geppert
🏛️ University of Stuttgart

In Time-Sensitive Networking (TSN), multicast communication improves bandwidth efficiency but exacerbates scheduling complexity due to port contention and queue resource constraints. Method: This paper proposes a fine-grained multicast tree partitioning approach that dynamically decomposes large multicast trees into smaller multicast or unicast subtrees, integrating time-triggered planning, adaptive tree-splitting algorithms, and multi-strategy scheduling to jointly optimize resource utilization and schedulability under heterogeneous topologies. Contribution/Results: To the best of our knowledge, this is the first systematic incorporation of multicast partitioning into TSN time-triggered flow scheduling. The method guarantees end-to-end latency bounds and achieves load balancing across switches. Experimental evaluation shows a 5–15% reduction in flow rejection rate and a 5–125% increase in throughput compared to an unpartitioned baseline, significantly improving flow admission ratio and scheduling feasibility in dynamic TSN environments.

Enhancing schedulability and throughput in dynamic network systemsOptimizing multicast communication in time-sensitive networksPartitioning multicast trees to improve bandwidth utilization

This work addresses the performance bottleneck of MPI_Alltoall under complex process layouts in high-dimensional torus topologies by proposing a zero-copy all-to-all communication algorithm. The approach decomposes the communication operation dimension by dimension, leveraging MPI Cartesian communicators and user-defined datatypes to implicitly perform data reordering without explicit local memory copies. A double-buffering mechanism and pre-cached dimensional communicators are introduced to enable flexible performance tuning. Experimental results demonstrate that the proposed method matches or even exceeds the performance of native MPI_Alltoall across various high-dimensional torus configurations, revealing significant optimization potential in current MPI implementations.

all-to-all communicationCartesian communicatorsMPI datatypes

Latest Papers

What's happening recently
View more

This work addresses the scalability bottleneck in MPI initialization caused by the reliance on the global communicator MPI_COMM_WORLD in traditional MPI implementations, particularly at exascale. By rearchitecting MPICH’s internal design, the authors present the first mainstream MPI implementation that fully decouples MPI_COMM_WORLD and introduces a compliant, true Sessions model aligned with the MPI-4 standard. The proposed approach employs explicit hierarchical process-set management and a scalable initialization protocol, substantially improving startup scalability. Experimental results demonstrate that the new mechanism efficiently supports exascale-class supercomputing systems, establishing a critical foundation for deploying MPI on next-generation ultra-large-scale platforms.

exascale systemsMPI SessionsMPI_COMM_WORLD

This study investigates the scalability and performance of process and thread schedulers under memory-intensive workloads in multi-core shared-memory systems, focusing on a 3D tensor row-sorting task. The authors design and evaluate several scheduling strategies: on the thread side, an AIMD-based adaptive chunking mechanism inspired by TCP congestion control is introduced, coupled with exponential weighted moving average to dynamically adjust concurrency; on the process side, a bounded prolific/collective model is employed alongside one-to-one, one-to-many, and many-to-many pipelined communication patterns to enable flexible task distribution. Experimental results on a 24-core x86-64 platform demonstrate that thread-level scheduling consistently outperforms process-level scheduling, with dynamic and guided strategies achieving the best performance, while the many-to-many pipeline exhibits superior scalability for large-scale tasks.

many-core systemsprocess-based schedulingscalability

This work addresses the inefficiency of MPI_Alltoallv in irregular communication scenarios, where repeated metadata processing degrades performance. For the first time, we introduce a persistent Remote Memory Access (RMA) mechanism into Alltoallv, decoupling initialization from execution to enable reuse of communication metadata and window state. We systematically evaluate the performance trade-offs between fence- and lock-based synchronization strategies. The proposed approach supports hierarchical scalability and demonstrates clear advantages at 448 processes with message sizes of 32 KB or larger. In large-message regimes, execution time improves from 2.49 seconds to 1.54 seconds, achieving up to a 44% speedup. This significantly enhances the scalability and practicality of irregular all-to-all communication in large-scale high-performance computing environments.

collective communicationhigh-performance computingMPI_Alltoallv

This work addresses the challenges of workflow task composition in high-throughput, petabyte-scale data processing environments, where resource heterogeneity and execution overhead significantly impact performance. The authors propose a hybrid task composition strategy that dynamically balances task independence against execution grouping, formulated within a multi-objective optimization framework to achieve Pareto-optimal trade-offs among throughput, I/O cost, and CPU efficiency. Leveraging workflow DAG modeling and high-dimensional parameter space simulation, the approach enables policy-driven automated synthesis of workflows. Experimental results demonstrate that the proposed strategy achieves up to a 3.8× improvement in throughput and reduces network overhead by as much as 14.9× compared to baseline methods, offering a scalable workflow synthesis framework for extreme-scale scientific computing.

extreme-scale data processingHigh-Throughput Computingresource utilization

Hot Scholars

FF

Fangcheng Fu

Shanghai Jiao Tong University
machine learningdeep learningMLSysdistributed computation
TH

Torsten Hoefler

Professor of Computer Science at ETH Zurich
High Performance ComputingDeep LearningNetworkingMessage Passing Interface
SS

Siddharth Singh

Research Scientist at Nvidia
High Performance ComputingArtificial Intelligence
HL

Haibin Lin

Bytedance
Machine Learning SystemsNatural Language Processing
KH

Kaibin Huang

Professor and Dept.Head, University of Hong Kong; NAI Fellow; IEEE Fellow; Highly Cited Researcher
Machine LearningMobile Edge ComputingWireless CommunicationsWireless Power Transfer