compute-communication overlap

Designs and evaluates methods and implementations that allow computation and communication to proceed concurrently rather than sequentially. Work includes creating nonblocking/asynchronous transfers, pipelining and buffering schemes, scheduling and runtime support, and metrics to measure and maximize overlap and its effect on performance.

compute-communicationoverlap

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Towards Bug-Free Distributed Go Programs

Jun 16, 2025
ZK
Zhengqun Koo
🏛️ National University of Singapore

Communication races in distributed Go programs—caused by violations of the happens-before relation—lead to erroneous message reception, including premature, missing, or partial message delivery. Method: We propose the first static verification framework for a Go subset that extends the happens-before ordering to both buffered and unbuffered channels, integrating formal channel semantics with abstractions of distributed execution traces. Contribution/Results: Our approach enables precise, sound static detection of communication races and fully verifies communication-race-freedom for representative message-driven distributed Go programs. By enforcing correct message ordering at compile time, it eliminates subtle runtime errors stemming from message reordering, thereby significantly enhancing the reliability of distributed systems. The framework is built upon rigorous modeling of Go’s concurrency primitives and supports automated, end-to-end verification without requiring program instrumentation or runtime monitoring.

Detecting communication races in distributed Go programsExtending happens-before analysis to Go channelsProving absence of races via static verification framework

This study systematically investigates the dual impact of computation-communication overlap on performance and energy efficiency in GPU-accelerated distributed deep learning. Addressing the challenges of increased computational latency, resource contention, and elevated power consumption under aggressive overlap, we build a multi-GPU distributed training testbed to evaluate overlap behavior across varying numerical precisions, specialized compute units (e.g., Tensor Cores), and power capping regimes. Results show that while overlap improves average throughput by 10.2%, it incurs up to 40% higher computational latency and an average 18.9% slowdown in computation; significant energy-efficiency trade-offs emerge under power and frequency constraints. We propose a novel “balanced overlap” strategy, demonstrating that hardware-specific characteristics—rather than overlap per se—are the primary determinants of its efficacy. Our findings provide empirical evidence and design principles for energy-aware distributed training systems.

Analyzes compute-communication overlap impact on GPU-accelerated distributed deep learning.Explores trade-offs between overlapping strategies, resource contention, and energy efficiency.Investigates hardware features' effect on performance and power in distributed training.

This work addresses the correctness challenges in implementing linearizable atomic registers in asynchronous message-passing systems, where precise real-time ordering of operations is unavailable. By combining equivalence and indistinguishability arguments with message-chain theory, the paper rigorously establishes that ensuring linearizability necessitates the formation of extensive message chains between operations of any type. This result formally characterizes, for the first time, the inherent communication overhead imposed by linearizability in asynchronous settings, thereby establishing a fundamental lower bound on the communication complexity required for its implementation. The findings provide a theoretical foundation for understanding the structural constraints and design costs associated with achieving linearizable semantics in distributed systems.

asynchronous systemsatomic registerscommunication requirements

This work addresses performance bottlenecks in existing servers during microsecond- to millisecond-scale hardware offloading, which often stem from context-switching overhead or busy-waiting. The authors propose a fine-grained offloading approach that requires no modification to server code, leveraging for the first time the server’s built-in suspend-resume concurrency mechanism and reframing offload scheduling as a routing problem. By injecting a fiber runtime via LD_PRELOAD, integrating native deferred response handling with executor submission, and employing page-protection techniques to ensure atomicity and safety, the method achieves significant speedups with only 22–138 lines of adapter code. Evaluated across ten widely used servers, it delivers 1.2–5.4× acceleration; notably, it attains a 17.3× speedup on unmodified thread-per-connection binaries and demonstrates both safety and efficacy in Redis.

computation offloadconcurrencyfine-grained

Convergence Sans Synchronization

Aug 09, 2025
AT
Arya Tanmay Gupta
🏛️ Michigan State University

Addressing the challenge of verifying convergence for parallel algorithms in asynchronous multiprocessor systems, this paper proposes a theoretical framework that eschews global synchronization mechanisms. Methodologically, it introduces a novel characterization of global convergence via partial orders among local state transition graphs—thereby avoiding the exponential construction and traversal of global state spaces inherent in conventional approaches. Integrating distributed state machine modeling with partial order theory, the framework enables the design of lock-free asynchronous algorithms whose convergence is formally provable and robust under arbitrary adversarial schedulers. Experimental evaluation demonstrates significantly improved convergence speed under asynchrony, alongside consistent efficiency and stability across both scheduler-sensitive and scheduler-insensitive scenarios. The results thus establish a rigorous theoretical foundation while maintaining practical applicability.

Develop fast parallel algorithms without costly synchronizationProve algorithm convergence in asynchrony without global state checksReduce complexity of verifying convergence via local partial order

Latest Papers

What's happening recently
View more

This work addresses the significant communication bottleneck in multi-GPU training caused by the serial execution of computation and communication. The authors propose a portable runtime mechanism that requires no modifications to vendor libraries or kernels. By dynamically controlling on-chip resource occupancy of compute kernels, elevating the scheduling priority of communication streams, and leveraging shared memory for compute footprint management and cross-GPU resource coordination, the approach effectively enables concurrent execution of computation and collective communication. Evaluated on NVIDIA A40, A100, H100, and AMD MI250X GPUs, the method reduces end-to-end training time by up to 25.5%.

communication overheadcomputation-communication overlapdistributed training

This study investigates concurrency progress conditions for linearizable shared objects that exploit commutativity-awareness in the asynchronous read-write shared-memory model. Addressing limitations of existing progress guarantees, the paper introduces a novel condition termed *conflict-obstruction-freedom*, which ensures that a process can complete its operation even when contending only with commutative operations. Leveraging formal tools from concurrent computation theory and linearizability semantics, the work provides the first rigorous formulation of this condition and proves that a universal construction satisfying it is impossible in the asynchronous read-write model. This impossibility result demonstrates that synchronization overhead remains unavoidable even when contention arises solely from conflicting—yet potentially commutative—operations, thereby establishing a fundamental limitation on progress guarantees in commutativity-aware concurrent systems.

asynchronous shared memorycommutativityconflict-obstruction-freedom

This work addresses the challenges of latency and scalability in designing efficient concurrent primitives under high write contention in shared-memory systems. It introduces a novel approach based on a contention-resolution algorithm that transforms contention-prone hardware primitives into higher-level concurrent objects within an approximately synchronous randomized scheduling model. For the first time, the study achieves composable, low-latency concurrent primitives against an adaptive adversary, and establishes a theoretical lower bound for the space–latency tradeoff. Using only O(1) read–write registers and a single compare-and-swap (CAS) register, the construction yields—with high probability—O(log P) latency for a variety of primitives, including read–write registers, CAS, load-linked/store-conditional (LL/SC), fetch-and-increment, bounded max registers, and counters.

concurrent primitivescontention resolutionlatency

This study investigates the scalability and performance of process and thread schedulers under memory-intensive workloads in multi-core shared-memory systems, focusing on a 3D tensor row-sorting task. The authors design and evaluate several scheduling strategies: on the thread side, an AIMD-based adaptive chunking mechanism inspired by TCP congestion control is introduced, coupled with exponential weighted moving average to dynamically adjust concurrency; on the process side, a bounded prolific/collective model is employed alongside one-to-one, one-to-many, and many-to-many pipelined communication patterns to enable flexible task distribution. Experimental results on a 24-core x86-64 platform demonstrate that thread-level scheduling consistently outperforms process-level scheduling, with dynamic and guided strategies achieving the best performance, while the many-to-many pipeline exhibits superior scalability for large-scale tasks.

many-core systemsprocess-based schedulingscalability

Hot Scholars

FF

Fangcheng Fu

Shanghai Jiao Tong University
machine learningdeep learningMLSysdistributed computation
DT

Dingwen Tao

Chinese Academy of Sciences, IEEE/ACM Senior Member
High Performance ComputingData ReductionDeep LearningSystems for ML
HX

Hongli Xu

University of Science and Technology of China
Software Defined NetworkCooperative CommunicationSensor Networks
LH

Liusheng Huang

Professor of Computer Science, University of Science and Technology of China
无线网络、信息安全
WH

Wenjing Huang

RAND Corporation
PsychometricsStructural Equation ModelingItem Response TheoryCyber Security