system latency optimization

Designs, implements, and evaluates low-latency systems and architectures, including networking and serving stacks, along with the profiling, monitoring, benchmarking, and tooling needed to measure and enforce latency budgets and cost–latency tradeoffs. Analyzes traces and profiles to find latency hotspots and applies optimizations across software, hardware, I/O, concurrency, serialization, caching, batching, and deployment/configuration to reduce average and tail latencies.

systemlatencyoptimization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$210K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses energy optimization for latency-sensitive network services under tail-latency SLA constraints. We propose a black-box, online co-tuning method that jointly optimizes interrupt coalescing (packet batching) and Dynamic Voltage and Frequency Scaling (DVFS), requiring no modifications to applications or the OS kernel. Our approach employs a generic, application- and system-agnostic Bayesian optimization controller that rapidly converges to the optimal energy-efficiency operating point under SLA constraints using minimal online probing. Key contributions include: (i) achieving up to 60% energy reduction while meeting tail-latency SLAs; (ii) revealing that specialized OS kernels improve energy efficiency by over 2× compared to general-purpose kernels; and (iii) demonstrating strong generalizability and stability across diverse hardware platforms. The method enables practical, deployment-ready energy-aware tuning for high-performance network services without sacrificing latency guarantees.

Achieve SLA targets efficientlyControl batching and processing rateOptimize energy-performance trade-offs

In the era of sub-millisecond networking, host-side latencies—such as those introduced by the kernel network stack and application scheduling—have become the dominant bottleneck for end-to-end low-latency performance, yet production environments lack effective means for continuous monitoring. This work proposes and implements netstacklat, the first system to enable low-overhead, continuous end-to-end latency monitoring within the Linux kernel network stack. By leveraging lightweight kernel probes and an efficient performance monitoring framework, netstacklat accurately captures the data path latency from the network interface card to the application across 144 diverse Nginx/Apache HTTP workloads, incurring less than 6% overhead even at tail latencies. The tool has been successfully deployed across Cloudflare’s global CDN infrastructure, demonstrating its scalability and practical utility in real-world production settings.

host latencylatency monitoringnetwork stack

Trace-based, time-resolved analysis of MPI application performance using standard metrics

Dec 01, 2025
KH
Kingshuk Haldar
🏛️ High Performance Computing Center Stuttgart | University of Stuttgart

Existing MPI performance analysis tools rely on time-aggregated metrics, which obscure transient bottlenecks. To address this, we propose a fine-grained, time-windowed trace analysis method that partitions execution traces into fixed or adaptive temporal windows and computes time-resolved metrics—including communication efficiency, load balance, and serialization overhead. Our approach integrates Paraver-based post-processing, critical path reconstruction, and event anomaly correction (e.g., clock skew compensation and unmatched MPI event reconciliation) to enable high-precision localization of transient bottlenecks. Evaluation on real-world applications (LaMEM, ls1-MarDyn) and synthetic benchmarks demonstrates that our method significantly improves both accuracy and scalability in identifying transient performance issues within large-scale traces, thereby overcoming the inherent limitations of global aggregation-based analysis.

Analyzes MPI application performance using time-resolved standard metricsIdentifies transient bottlenecks hidden by time-aggregated metrics in toolsProcesses execution traces robustly despite anomalies and large sizes

How long can you sleep? Idle Time System Inefficiencies and Opportunities

Oct 08, 2025
GA
Georgia Antoniou
🏛️ University of Cyprus | Rivos Inc.

To address low utilization of deep idle states in latency-sensitive applications, this paper identifies a significant gap between theoretically available idle opportunities and their actual exploitation—caused by inaccurate idle scheduling decisions and non-negligible deep-sleep transition latency. We propose a queueing-theoretic modeling framework that integrates M/M/1, c×M/M/1, and M/M/c models, calibrated with real-world server workload traces, to quantify system-level idle potential under diverse configurations. For the first time, we systematically identify numerous untriggered deep-idle entry opportunities and develop a scalable methodology for idle-efficiency evaluation. The framework provides quantifiable, early-stage guidance for hardware–OS co-design, enabling energy-efficiency optimization and supporting system-level power management strategies that explicitly balance latency constraints and energy savings.

Identifying inefficiencies in entering deep idle power statesModeling idle time distribution in latency-critical server systemsProviding early-stage design exploration for server configurations

Performance Models for a Two-tiered Storage System

Mar 12, 2025
AS
Aparna Sasidharan
🏛️ IIT | Sandia National Lab | Oak Ridge National Lab

To address inefficient data migration and inaccurate performance prediction in heterogeneous storage systems (NVMe cache + HDD backend), this paper designs and implements a distributed two-tier storage system. We propose an online reinforcement learning–based dynamic data tiering scheduling algorithm and develop an end-to-end performance model integrating queuing network theory with fine-grained device behavior modeling. Our key contribution is the first scalable, fine-grained device behavior modeling method tailored for heterogeneous storage—enabling adaptive tiering management and precise performance prediction under high-concurrency I/O workloads in multi-core clusters. Experimental evaluation on multi-node clusters demonstrates an average model prediction error of less than 8%, a 27% improvement in I/O throughput, and a 34% reduction in average access latency. The framework provides a reusable modeling and optimization foundation for two-tier storage systems.

Design and analyze a two-tiered storage systemDevelop online learning for data tier managementEvaluate performance using queuing and behavioral models

Latest Papers

What's happening recently
View more

This work challenges the conventional focus on core utilization in resource management, which often overlooks the practical performance constraints imposed by power and thermal limits in modern multicore processors. Instead, it proposes a new paradigm centered on power budgeting, elevating idle-core waiting strategies to first-class design considerations. Rather than aggressively reclaiming idle cores—a practice that frequently overestimates benefits and incurs substantial scheduling overhead—the approach leverages efficient waiting mechanisms to release redistributable compute capacity. Empirical analysis on AMD EPYC platforms, accounting for processor topology, idle duration, and waiting policies, demonstrates that such strategies achieve a superior trade-off between energy efficiency and performance, offering greater practical advantages in real-world systems.

idle coremulticore processorspolling efficiency

This paper identifies a critical gap in high-performance data transfer research: an overemphasis on network bandwidth while neglecting end-to-end bottlenecks—including latency, TCP congestion control, host CPU limitations, and virtualization—leading to severe discrepancies between benchmark results and real-world production performance. To address this, the authors propose a hardware–software co-design paradigm and develop a latency-programmable testbed. Leveraging high-fidelity wide-area network (WAN) modeling and cross-continental 100 Gbps measurements (Switzerland–California), they systematically isolate key constraints at the network edge and host side. Results demonstrate that primary bottlenecks reside predominantly at the network edge—not the core—and that stable, predictable data movement is achieved across 1–100+ Gbps. This significantly enhances performance fidelity in complex, heterogeneous environments.

Examines host-side factors like CPU and virtualization impacting workflows.Investigates bottlenecks beyond network bandwidth in data movement.Proposes holistic hardware-software co-design for consistent performance.

This work addresses the challenge of meeting latency constraints in multi-model large language model (LLM) serving, where the tight coupling between request routing and resource allocation renders traditional approaches ineffective. To tackle this issue, the paper presents the first joint optimization framework that simultaneously models both decisions. The approach constructs a deployment-aware latency model based on empirical system measurements and leverages a dual pricing mechanism to solve the constrained optimization problem under latency service-level objectives (SLOs). Experimental results demonstrate that, on the same GPU cluster, varying resource allocations can lead to up to an 87% difference in output quality, highlighting the critical importance of co-optimizing routing and resource provisioning. The proposed framework effectively enhances service quality while rigorously satisfying latency requirements.

GPU clusterslatency SLOmulti-model LLM serving

Hot Scholars

KH

Kaibin Huang

Professor and Dept.Head, University of Hong Kong; NAI Fellow; IEEE Fellow; Highly Cited Researcher
Machine LearningMobile Edge ComputingWireless CommunicationsWireless Power Transfer
YH

Yao Hu

浙江大学
Machine Learning
XC

Xianhao Chen

Assistant Professor, The University of Hong Kong
Wireless networksmobile edge computingedge AIdistributed learning
MG

Minyi Guo

IEEE Fellow, Chair Professor, Shanghai Jiao Tong University
Parallel ComputingCompiler OptimizationCloud ComputingNetworking
LZ

Ligeng Zhu

Nvidia
Machine LearningEfficient Deep Learning