latency hiding

Designs and implements runtime schedules, pipelined execution, and data-transfer/buffering schemes that overlap communication, I/O, and computation so that communication or memory delays do not stall forward progress. Builds or analyzes mechanisms—such as asynchronous transfers, double buffering, prefetching, and state-transfer protocols—to maintain throughput and responsiveness under bursty or long-context request patterns.

latencyhiding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.09
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the correctness challenges in implementing linearizable atomic registers in asynchronous message-passing systems, where precise real-time ordering of operations is unavailable. By combining equivalence and indistinguishability arguments with message-chain theory, the paper rigorously establishes that ensuring linearizability necessitates the formation of extensive message chains between operations of any type. This result formally characterizes, for the first time, the inherent communication overhead imposed by linearizability in asynchronous settings, thereby establishing a fundamental lower bound on the communication complexity required for its implementation. The findings provide a theoretical foundation for understanding the structural constraints and design costs associated with achieving linearizable semantics in distributed systems.

asynchronous systemsatomic registerscommunication requirements

This work addresses performance bottlenecks in existing servers during microsecond- to millisecond-scale hardware offloading, which often stem from context-switching overhead or busy-waiting. The authors propose a fine-grained offloading approach that requires no modification to server code, leveraging for the first time the server’s built-in suspend-resume concurrency mechanism and reframing offload scheduling as a routing problem. By injecting a fiber runtime via LD_PRELOAD, integrating native deferred response handling with executor submission, and employing page-protection techniques to ensure atomicity and safety, the method achieves significant speedups with only 22–138 lines of adapter code. Evaluated across ten widely used servers, it delivers 1.2–5.4× acceleration; notably, it attains a 17.3× speedup on unmodified thread-per-connection binaries and demonstrates both safety and efficacy in Redis.

computation offloadconcurrencyfine-grained

Work in Progress: Middleware-Transparent Callback Enforcement in Commoditized Component-Oriented Real-time Systems

May 10, 2025
TI
Takahiro Ishikawa-Aso
🏛️ The University of Tokyo | TIER IV Incorporated | Saitama University

To address the high scheduling overhead caused by nested scheduling (OS threads + middleware Executors) in commercial real-time systems such as ROS 2, this paper proposes a lightweight, real-time scheduling paradigm—“one-to-one binding of callbacks to OS threads”—which bypasses the middleware scheduling layer and enables native OS-level scheduling control at the callback granularity. Our key contributions are: (1) the first middleware-transparent callback scheduling model, eliminating nested scheduling complexity; and (2) CallbackIsolatedExecutor, a novel executor that supports direct configuration of kernel-level parameters—including SCHED_FIFO, priority, and CPU affinity. Experimental results show that, compared to MultiThreadedExecutor, our approach significantly reduces context switches, user-to-kernel transitions, and memory overhead. Against SingleThreadedExecutor, inter-process and intra-process communication latencies remain consistently at 1.4× and 5×, respectively—achieving a balanced trade-off between real-time determinism and schedulability control.

Avoiding nested scheduling in ROS 2 real-time researchEnforcing direct OS scheduling on callbacks in real-time systemsReducing middleware layer costs in component-oriented systems

Shared state profoundly influences the performance and fault tolerance of stream processing, service-oriented, and continual learning systems, yet existing approaches often treat access control, hardware-aware execution, memory management, and long-term evolution in isolation. This work reframes state management as a runtime control problem and introduces a contract-driven blueprint centered on state objects, control planes, coupling paths, evaluation boundaries, and pending contracts. Building upon this foundation, we develop a unified analytical framework encompassing state-access scheduling, state-aware execution, and state evolution reuse. Through systematic scheduling, runtime control, and cross-layer coupling analysis, our approach identifies critical anti-patterns and advances a perturbation-aware evaluation paradigm, thereby establishing both theoretical foundations and practical design guidelines for state control in distributed systems.

distributed systemsparallel systemsruntime control

Indirect Coflow Scheduling

Nov 16, 2025
AL
Alexander Lindermayr
🏛️ Simons Institute for the Theory of Computing | UC Berkeley | University of Pittsburgh | Arizona State University | Khoury College of Computer Sciences | Northeastern University

While large-scale data transfers in reconfigurable networks have been extensively studied, indirect cooperative flow scheduling for small-scale requests—relative to single-round transmission capacity—has long been overlooked, leading to prolonged completion times and low resource utilization. Method: This paper presents the first systematic modeling and optimization of cooperative flow scheduling under this scenario. We propose a combinatorial optimization framework integrating fractional matching and indirect routing: fractional matching enables fine-grained bandwidth allocation, while multi-hop indirect paths relax direct-connectivity constraints, supporting demand-driven elastic scheduling. Building upon theoretical schedulability analysis, we design an efficient heuristic algorithm. Results: Experiments demonstrate that, in small-scale data transfer scenarios, our approach reduces average flow completion time by 32.7% and improves link resource utilization by 41.5% over state-of-the-art methods, significantly enhancing scheduling efficiency and network adaptability for lightweight traffic.

Comparing indirect routing and fractional matchings for efficiencyDesigning algorithms optimized for small demands versus large transfersScheduling coflows for small data transfers in reconfigurable networks

Latest Papers

What's happening recently
View more

This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.

Autonomous AgentsHigh Performance ComputingJob Specification Translation

Minimize Your Critical Path with Combine-and-Exchange Locks

Nov 12, 2025
SK
Simon König
🏛️ University of Stuttgart

Existing user-space coroutine/fiber synchronization mechanisms implicitly assume kernel scheduling, introducing unnecessary latency on critical paths and limiting high-concurrency throughput. This paper proposes Combine-and-Exchange Scheduling (CES), a novel synchronization paradigm for purely user-space cooperative scheduling. CES eliminates cross-thread overhead by retaining critical sections on the same thread during lock contention, while dynamically redistributing parallelizable tasks to idle threads. Crucially, it co-designs user-space synchronization primitives with the scheduler to fully bypass kernel intervention. Experimental evaluation demonstrates that CES achieves up to 3× higher throughput on application-level benchmarks and up to 8× speedup on microbenchmarks—significantly outperforming state-of-the-art user-space synchronization approaches.

Improving throughput for coroutine-based parallel applicationsOptimizing scheduling for contended critical sections across threadsReducing critical path delays in userspace synchronization primitives

This work addresses the limitations of traditional causal delivery approaches in large-scale distributed systems, which suffer from high metadata overhead, limited throughput scalability, or dependence on specific network topologies. The authors propose a topology-agnostic hybrid buffering mechanism that integrates a Sender Permission to Send (SPS) policy at the sender side with FIFO buffering at the receiver side. This approach achieves, for the first time, constant per-message metadata and amortized constant computational overhead while preserving causal consistency. By transcending the constraints of purely sender-side or receiver-side solutions, the method significantly enhances throughput and scalability, establishing a new paradigm for efficient causal message delivery in distributed systems.

causal deliverymetadata overheadsender buffering

As per-port bandwidth in data center switches continues to increase, traditional buffer-sharing strategies suffer from degraded performance and high complexity. This work proposes BShare, a lightweight, queueing-delay-based buffer-sharing mechanism that integrates buffer management with active queue management using only a single configurable parameter. By dynamically monitoring queueing delay to adjust buffer allocation in real time, BShare significantly simplifies policy design while maintaining compatibility with advanced transport protocols such as PowerTCP. Simulation results demonstrate that under burst-intensive workloads, BShare reduces flow completion time (FCT) by up to 45.07% compared to the ABM scheme.

buffer managementbuffer sharingdatacenter switches

Hot Scholars

SH

Sitao Huang

Assistant Professor of EECS, University of California Irvine
Hardware AccelerationHigh-Level SynthesisFPGAParallel Computing
YQ

Ye Qiao

Ph.D. Candidate, University of California, Irvine
Machine LearningComputer ArchitectureComputer VisionEdge Computing
WW

Wei Wang

The Hong Kong University of Science and Technology
Cloud ComputingMachine Learning SystemsBig Data SystemsComputer Networking