latency

Designs, measures, and optimizes the response-time and delay characteristics of systems and components, including mean and tail latency, jitter, and latency distributions. Analyzes sources of delay and trade-offs with throughput or consistency, and implements techniques (e.g., batching, pipelining, caching, concurrency control, scheduling, and network/configuration tuning) to meet latency targets and SLOs.

latency

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
2.65
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$212K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses energy optimization for latency-sensitive network services under tail-latency SLA constraints. We propose a black-box, online co-tuning method that jointly optimizes interrupt coalescing (packet batching) and Dynamic Voltage and Frequency Scaling (DVFS), requiring no modifications to applications or the OS kernel. Our approach employs a generic, application- and system-agnostic Bayesian optimization controller that rapidly converges to the optimal energy-efficiency operating point under SLA constraints using minimal online probing. Key contributions include: (i) achieving up to 60% energy reduction while meeting tail-latency SLAs; (ii) revealing that specialized OS kernels improve energy efficiency by over 2× compared to general-purpose kernels; and (iii) demonstrating strong generalizability and stability across diverse hardware platforms. The method enables practical, deployment-ready energy-aware tuning for high-performance network services without sacrificing latency guarantees.

Achieve SLA targets efficientlyControl batching and processing rateOptimize energy-performance trade-offs

Characterization of latency and jitter in TSN emulation

Jun 02, 2025
AG
Alex Gracia
🏛️ Universidad de Zaragoza | Intel Corporation | CINVESTAV

Existing software simulation of Time-Sensitive Networking (TSN) suffers from insufficient accuracy in measuring bridge delay and jitter, undermining the fidelity and reproducibility of TSN emulation. Method: This paper introduces the first systematic timestamping methodology for TSN simulation on Linux/Mininet, rigorously evaluating four timestamping mechanisms—including SO_TIMESTAMPING—under TSN traffic shaped by Credit-Based Shaping (CBS) and Asynchronous Traffic Shaping (ATS). Leveraging configurable Mininet topologies, the approach integrates scheduling solution generation, deployment validation, and cross-platform optimization—supporting both Intel Time-Coordinated Computing (TCC)-enabled and -disabled modes on industrial PCs and workstations. Contribution/Results: The framework achieves sub-microsecond bridge delay characterization and, for the first time, experimentally validates end-to-end deterministic guarantees on real hardware. It overcomes critical bottlenecks in clock synchronization precision and scheduling fidelity, significantly enhancing the trustworthiness and reproducibility of TSN simulation.

Characterizing latency and jitter in TSN emulation environmentsEvaluating timestamping methods for TSN network traffic profilingSolving TSN scheduling challenges in software-based emulation

In the era of sub-millisecond networking, host-side latencies—such as those introduced by the kernel network stack and application scheduling—have become the dominant bottleneck for end-to-end low-latency performance, yet production environments lack effective means for continuous monitoring. This work proposes and implements netstacklat, the first system to enable low-overhead, continuous end-to-end latency monitoring within the Linux kernel network stack. By leveraging lightweight kernel probes and an efficient performance monitoring framework, netstacklat accurately captures the data path latency from the network interface card to the application across 144 diverse Nginx/Apache HTTP workloads, incurring less than 6% overhead even at tail latencies. The tool has been successfully deployed across Cloudflare’s global CDN infrastructure, demonstrating its scalability and practical utility in real-world production settings.

host latencylatency monitoringnetwork stack

Characterizing the Impact of Active Queue Management on Speed Test Measurements

Nov 24, 2025
SR
Siddhant Ray
🏛️ University of Chicago | Cal Poly | ENS Lyon

Existing speed measurement tools focus on peak throughput and poorly reflect users’ perceived responsiveness; emerging metrics such as “latency under load” show promise but their sensitivity to Active Queue Management (AQM) configurations remains unclear. Method: We empirically evaluate three mainstream AQM schemes—CoDel, FQ-CoDel, and SFQ—in a controlled network environment, systematically analyzing their impact on throughput and latency distributions, particularly latency under load. Results: AQM significantly alters speed test outcomes, with distinct latency-throughput trade-offs observed across algorithms under high load. Current measurement platforms, if uncalibrated for AQM, yield misleading latency estimates, undermining the reliability of policy and regulatory decisions. This study is the first to quantitatively characterize the structural impact of AQM on emerging speed metrics, providing critical empirical evidence to inform standardization of measurement tools and evidence-based network governance.

Calibrating speed tests for accurate policy guidanceComparing throughput variance across different AQM schemesUnderstanding AQM's impact on speed test latency metrics

Latest Papers

What's happening recently
View more

This work addresses load balancing in parallel infinite-server queues under action delays by explicitly incorporating delay into the state representation for the first time. The authors model the delay process using an Erlang phase-type structure, thereby constructing a finite-dimensional Markov jump system and employing ordinary differential equations to explicitly track tasks in transit. Exploiting symmetry between two servers, the system dynamics are reduced to a single mode capturing load imbalance. Theoretical analysis establishes equivalence between this formulation and the delayed-information model in their linearized dynamics, overcoming limitations of traditional delay-differential-equation approaches. The study derives the characteristic equation governing imbalance dynamics for an arbitrary number of phases, and numerical experiments confirm the accuracy of the fluid approximation while quantifying the effects of phase count, routing sensitivity, and mean delay on transient response.

action delayinfinite-server queuesload balancing

In cloud-edge collaborative environments, microservices struggle to simultaneously address computational bottlenecks and end-to-end latency due to dynamic factors such as heterogeneous nodes, time-varying network delays, non-stationary traffic, and mixed request patterns, often leading to violations of service-level objectives (SLOs). To tackle this challenge, this work proposes ADASCALE, a novel framework that uniquely integrates root operation identification, SLO-aware scaling, and a demand-weighted latency-minimizing placement strategy. ADASCALE employs a dual-loop control mechanism—combining reactive and steady-state control—to jointly optimize microservice scaling and cross-node scheduling. Leveraging distributed tracing and service mesh metrics, ADASCALE demonstrates significant improvements over NetMARKS_Scale on the DeathStarBench benchmark, reducing average response time by 1.93× and increasing throughput by 2.16× while consistently meeting SLO requirements.

autoscalingcloud-edgemicroservices

This work addresses the inefficiencies in multi-cluster cloud data warehouses, where static or over-provisioned resource allocation often leads to excessive costs and violations of latency service-level objectives (SLOs). To tackle this, the authors propose AutoSLO, a novel framework that integrates historical workload forecasting, real-time responsive scaling, and concurrency-aware query routing to jointly optimize resource cost and SLO compliance across multiple timescales. AutoSLO features a three-tier architecture comprising a policy tuner leveraging one-day historical data, an SLO-aware autoscaler, and an online query router. Experimental evaluation on Redbench demonstrates that AutoSLO reduces average costs by 26.4%, with the query router and autoscaler lowering SLO violation rates by 47.8% and 93.7%, respectively, while the policy tuner alone achieves a 44.6% reduction in violations using only a single day of historical data.

cloud data warehousescluster scalinglatency SLO

This work addresses a critical limitation in existing latency models for strict-priority scheduling, which neglect the impact of the transmit ring buffer (TXR) at switch egress ports, leading to inaccurate delay estimates for high-priority packets. For the first time, this study incorporates TXR into the end-to-end delay analysis framework by proposing a method to measure TXR size and integrating deterministic and stochastic modeling, network measurements, and hardware behavior characterization. The resulting model significantly improves the accuracy of both worst-case delay bounds and delay distribution predictions for high-priority traffic. Validated across multiple commercial off-the-shelf switches, the approach effectively bridges the gap between theoretical models and real hardware behavior, thereby enhancing the reliability of real-time network system design.

Latency ModelingPacket DelayStrict Priority