Score
Designs and evaluates methods and implementations that allow computation and communication to proceed concurrently rather than sequentially. Work includes creating nonblocking/asynchronous transfers, pipelining and buffering schemes, scheduling and runtime support, and metrics to measure and maximize overlap and its effect on performance.
Communication races in distributed Go programs—caused by violations of the happens-before relation—lead to erroneous message reception, including premature, missing, or partial message delivery. Method: We propose the first static verification framework for a Go subset that extends the happens-before ordering to both buffered and unbuffered channels, integrating formal channel semantics with abstractions of distributed execution traces. Contribution/Results: Our approach enables precise, sound static detection of communication races and fully verifies communication-race-freedom for representative message-driven distributed Go programs. By enforcing correct message ordering at compile time, it eliminates subtle runtime errors stemming from message reordering, thereby significantly enhancing the reliability of distributed systems. The framework is built upon rigorous modeling of Go’s concurrency primitives and supports automated, end-to-end verification without requiring program instrumentation or runtime monitoring.
This study systematically investigates the dual impact of computation-communication overlap on performance and energy efficiency in GPU-accelerated distributed deep learning. Addressing the challenges of increased computational latency, resource contention, and elevated power consumption under aggressive overlap, we build a multi-GPU distributed training testbed to evaluate overlap behavior across varying numerical precisions, specialized compute units (e.g., Tensor Cores), and power capping regimes. Results show that while overlap improves average throughput by 10.2%, it incurs up to 40% higher computational latency and an average 18.9% slowdown in computation; significant energy-efficiency trade-offs emerge under power and frequency constraints. We propose a novel “balanced overlap” strategy, demonstrating that hardware-specific characteristics—rather than overlap per se—are the primary determinants of its efficacy. Our findings provide empirical evidence and design principles for energy-aware distributed training systems.
This work addresses the correctness challenges in implementing linearizable atomic registers in asynchronous message-passing systems, where precise real-time ordering of operations is unavailable. By combining equivalence and indistinguishability arguments with message-chain theory, the paper rigorously establishes that ensuring linearizability necessitates the formation of extensive message chains between operations of any type. This result formally characterizes, for the first time, the inherent communication overhead imposed by linearizability in asynchronous settings, thereby establishing a fundamental lower bound on the communication complexity required for its implementation. The findings provide a theoretical foundation for understanding the structural constraints and design costs associated with achieving linearizable semantics in distributed systems.
This work addresses performance bottlenecks in existing servers during microsecond- to millisecond-scale hardware offloading, which often stem from context-switching overhead or busy-waiting. The authors propose a fine-grained offloading approach that requires no modification to server code, leveraging for the first time the server’s built-in suspend-resume concurrency mechanism and reframing offload scheduling as a routing problem. By injecting a fiber runtime via LD_PRELOAD, integrating native deferred response handling with executor submission, and employing page-protection techniques to ensure atomicity and safety, the method achieves significant speedups with only 22–138 lines of adapter code. Evaluated across ten widely used servers, it delivers 1.2–5.4× acceleration; notably, it attains a 17.3× speedup on unmodified thread-per-connection binaries and demonstrates both safety and efficacy in Redis.
Addressing the challenge of verifying convergence for parallel algorithms in asynchronous multiprocessor systems, this paper proposes a theoretical framework that eschews global synchronization mechanisms. Methodologically, it introduces a novel characterization of global convergence via partial orders among local state transition graphs—thereby avoiding the exponential construction and traversal of global state spaces inherent in conventional approaches. Integrating distributed state machine modeling with partial order theory, the framework enables the design of lock-free asynchronous algorithms whose convergence is formally provable and robust under arbitrary adversarial schedulers. Experimental evaluation demonstrates significantly improved convergence speed under asynchrony, alongside consistent efficiency and stability across both scheduler-sensitive and scheduler-insensitive scenarios. The results thus establish a rigorous theoretical foundation while maintaining practical applicability.
This work addresses the significant communication bottleneck in multi-GPU training caused by the serial execution of computation and communication. The authors propose a portable runtime mechanism that requires no modifications to vendor libraries or kernels. By dynamically controlling on-chip resource occupancy of compute kernels, elevating the scheduling priority of communication streams, and leveraging shared memory for compute footprint management and cross-GPU resource coordination, the approach effectively enables concurrent execution of computation and collective communication. Evaluated on NVIDIA A40, A100, H100, and AMD MI250X GPUs, the method reduces end-to-end training time by up to 25.5%.
This study investigates concurrency progress conditions for linearizable shared objects that exploit commutativity-awareness in the asynchronous read-write shared-memory model. Addressing limitations of existing progress guarantees, the paper introduces a novel condition termed *conflict-obstruction-freedom*, which ensures that a process can complete its operation even when contending only with commutative operations. Leveraging formal tools from concurrent computation theory and linearizability semantics, the work provides the first rigorous formulation of this condition and proves that a universal construction satisfying it is impossible in the asynchronous read-write model. This impossibility result demonstrates that synchronization overhead remains unavoidable even when contention arises solely from conflicting—yet potentially commutative—operations, thereby establishing a fundamental limitation on progress guarantees in commutativity-aware concurrent systems.
This work addresses the challenges of latency and scalability in designing efficient concurrent primitives under high write contention in shared-memory systems. It introduces a novel approach based on a contention-resolution algorithm that transforms contention-prone hardware primitives into higher-level concurrent objects within an approximately synchronous randomized scheduling model. For the first time, the study achieves composable, low-latency concurrent primitives against an adaptive adversary, and establishes a theoretical lower bound for the space–latency tradeoff. Using only O(1) read–write registers and a single compare-and-swap (CAS) register, the construction yields—with high probability—O(log P) latency for a variety of primitives, including read–write registers, CAS, load-linked/store-conditional (LL/SC), fetch-and-increment, bounded max registers, and counters.
This study investigates the scalability and performance of process and thread schedulers under memory-intensive workloads in multi-core shared-memory systems, focusing on a 3D tensor row-sorting task. The authors design and evaluate several scheduling strategies: on the thread side, an AIMD-based adaptive chunking mechanism inspired by TCP congestion control is introduced, coupled with exponential weighted moving average to dynamically adjust concurrency; on the process side, a bounded prolific/collective model is employed alongside one-to-one, one-to-many, and many-to-many pipelined communication patterns to enable flexible task distribution. Experimental results on a 24-core x86-64 platform demonstrate that thread-level scheduling consistently outperforms process-level scheduling, with dynamic and guided strategies achieving the best performance, while the many-to-many pipeline exhibits superior scalability for large-scale tasks.