real-time latency characterization

Designs and implements measurement, profiling, and benchmarking artifacts (testbeds, probes, trace collection, benchmark suites) to measure and characterize latency and throughput properties of systems and services in real-time or near-real-time settings, including end-to-end, network, and tail-latency behavior. Builds analytical and statistical models and tooling to estimate and predict latency distributions, performs latency and memory/throughput profiling, and conducts latency‑accuracy and other tradeoff analyses to validate real-time constraints and guide latency-focused optimization.

real-timelatencycharacterization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.72
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$222K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Characterizing the Impact of Active Queue Management on Speed Test Measurements

Nov 24, 2025
SR
Siddhant Ray
🏛️ University of Chicago | Cal Poly | ENS Lyon

Existing speed measurement tools focus on peak throughput and poorly reflect users’ perceived responsiveness; emerging metrics such as “latency under load” show promise but their sensitivity to Active Queue Management (AQM) configurations remains unclear. Method: We empirically evaluate three mainstream AQM schemes—CoDel, FQ-CoDel, and SFQ—in a controlled network environment, systematically analyzing their impact on throughput and latency distributions, particularly latency under load. Results: AQM significantly alters speed test outcomes, with distinct latency-throughput trade-offs observed across algorithms under high load. Current measurement platforms, if uncalibrated for AQM, yield misleading latency estimates, undermining the reliability of policy and regulatory decisions. This study is the first to quantitatively characterize the structural impact of AQM on emerging speed metrics, providing critical empirical evidence to inform standardization of measurement tools and evidence-based network governance.

Calibrating speed tests for accurate policy guidanceComparing throughput variance across different AQM schemesUnderstanding AQM's impact on speed test latency metrics

In the era of sub-millisecond networking, host-side latencies—such as those introduced by the kernel network stack and application scheduling—have become the dominant bottleneck for end-to-end low-latency performance, yet production environments lack effective means for continuous monitoring. This work proposes and implements netstacklat, the first system to enable low-overhead, continuous end-to-end latency monitoring within the Linux kernel network stack. By leveraging lightweight kernel probes and an efficient performance monitoring framework, netstacklat accurately captures the data path latency from the network interface card to the application across 144 diverse Nginx/Apache HTTP workloads, incurring less than 6% overhead even at tail latencies. The tool has been successfully deployed across Cloudflare’s global CDN infrastructure, demonstrating its scalability and practical utility in real-world production settings.

host latencylatency monitoringnetwork stack

Trace-based, time-resolved analysis of MPI application performance using standard metrics

Dec 01, 2025
KH
Kingshuk Haldar
🏛️ High Performance Computing Center Stuttgart | University of Stuttgart

Existing MPI performance analysis tools rely on time-aggregated metrics, which obscure transient bottlenecks. To address this, we propose a fine-grained, time-windowed trace analysis method that partitions execution traces into fixed or adaptive temporal windows and computes time-resolved metrics—including communication efficiency, load balance, and serialization overhead. Our approach integrates Paraver-based post-processing, critical path reconstruction, and event anomaly correction (e.g., clock skew compensation and unmatched MPI event reconciliation) to enable high-precision localization of transient bottlenecks. Evaluation on real-world applications (LaMEM, ls1-MarDyn) and synthetic benchmarks demonstrates that our method significantly improves both accuracy and scalability in identifying transient performance issues within large-scale traces, thereby overcoming the inherent limitations of global aggregation-based analysis.

Analyzes MPI application performance using time-resolved standard metricsIdentifies transient bottlenecks hidden by time-aggregated metrics in toolsProcesses execution traces robustly despite anomalies and large sizes

Characterization of latency and jitter in TSN emulation

Jun 02, 2025
AG
Alex Gracia
🏛️ Universidad de Zaragoza | Intel Corporation | CINVESTAV

Existing software simulation of Time-Sensitive Networking (TSN) suffers from insufficient accuracy in measuring bridge delay and jitter, undermining the fidelity and reproducibility of TSN emulation. Method: This paper introduces the first systematic timestamping methodology for TSN simulation on Linux/Mininet, rigorously evaluating four timestamping mechanisms—including SO_TIMESTAMPING—under TSN traffic shaped by Credit-Based Shaping (CBS) and Asynchronous Traffic Shaping (ATS). Leveraging configurable Mininet topologies, the approach integrates scheduling solution generation, deployment validation, and cross-platform optimization—supporting both Intel Time-Coordinated Computing (TCC)-enabled and -disabled modes on industrial PCs and workstations. Contribution/Results: The framework achieves sub-microsecond bridge delay characterization and, for the first time, experimentally validates end-to-end deterministic guarantees on real hardware. It overcomes critical bottlenecks in clock synchronization precision and scheduling fidelity, significantly enhancing the trustworthiness and reproducibility of TSN simulation.

Characterizing latency and jitter in TSN emulation environmentsEvaluating timestamping methods for TSN network traffic profilingSolving TSN scheduling challenges in software-based emulation

Measuring Round-Trip Response Latencies Under Asymmetric Routing

May 20, 2025
BV
Bhavana Vannarth Shobhana
🏛️ Rutgers University | The Open University of Israel

This paper addresses the challenge of passively measuring end-to-end response latency under transport-layer encryption (e.g., TLS/QUIC), where application-layer headers and client instrumentation are unavailable. We propose PIRATE, a passive latency estimation algorithm that relies solely on observable client→server traffic. Its core innovation is the first use of causally linked request-pair time differences—without client-side instrumentation or plaintext headers—to accurately proxy application-layer round-trip latency, inherently supporting asymmetric routing. The method integrates causal request-pair detection, passive temporal modeling, and DSR-aware load-balancing coordination. Evaluated on real-world web services, PIRATE achieves ≤1% estimation error for client-side latency. When deployed at Layer-4 load balancers, it reduces tail latency by 37%, significantly enhancing bottleneck identification, adaptive scheduling, and attack mitigation capabilities.

Improve load balancing by reducing tail latencies significantlyMeasure response latencies under asymmetric routing conditionsPassively estimate client-side latency without client instrumentation

Latest Papers

What's happening recently
View more

This work addresses the lack of existing tools capable of continuous performance validation and regression detection across entire datacenter clusters. The authors propose the first cluster-wide continuous benchmarking framework that supports unified scheduling, enabling simultaneous task distribution to all nodes and systematic collection of multidimensional performance metrics—spanning CPU, GPU, memory, interconnects, I/O, power consumption, frequency, and temperature—across both space and time. This framework facilitates performance regression detection under software and hardware changes as well as analysis of hardware variability. Experiments on the NHR@FAU cluster reveal intra-node performance variations below 1% among identically configured nodes, while inter-node differences reach up to 5%. The study further uncovers, for the first time, significant disparities in the performance–power relationship between air-cooled and liquid-cooled nodes.

cluster-wide benchmarkingcontinuous testingdata center validation

This study addresses the limited sensitivity of traditional cloud service performance regression detection, which is often hindered by I/O fluctuations and infrastructure changes. The authors propose a novel paradigm termed “Duet Instrumentation,” which uniquely integrates large language model (LLM)-driven code change analysis with synchronized dual-version benchmarking. By leveraging an LLM to precisely identify performance-relevant changes between consecutive versions, the method dynamically instruments only those critical code regions, achieving high-sensitivity regression detection with low overhead. Evaluated in real-world environments, the approach attains a precision of 58%, recall of 93%, and specificity of 71%, effectively detecting performance regressions as subtle as one-fifth the severity detectable by conventional methods.

application benchmarkscloud service benchmarkingmicrobenchmarks

This work addresses behavioral inconsistencies between Linux kernel implementations and userspace simulators of the L4S (Low Latency, Low Loss, Scalable throughput) mechanism, which hinder experimental reproducibility and parameter portability. We present the first scalable implementation of the DualPI2 active queue management algorithm in Mahimahi and conduct a systematic comparison against its kernel counterpart across diverse traffic patterns and network conditions. Through comprehensive behavioral characterization and parameter sensitivity analysis, we identify the bandwidth-delay product (BDP) as a critical factor governing cross-platform discrepancies. Our findings reveal specific parameter configurations that improve alignment under low-BDP scenarios, while also exposing persistent structural deviations under high load. This study provides both a practical simulation tool and empirical guidance for accurate L4S experimentation and deployment.

active queue managementbehavioral characterizationcross-platform analysis

This work proposes a cross-layer, interpretable performance diagnosis method to address the challenge of detecting subtle radio-layer dynamic anomalies in O-RAN systems when end-to-end latency appears stable. Leveraging real-world measurements across multiple distances and user equipment (UE) types, the approach jointly analyzes application-layer tail latency—such as the 95th percentile—with radio-layer metrics including scheduling behavior, modulation and coding scheme (MCS), block error rate (BLER), and signal quality to construct lightweight “degradation flags.” The method enables non-intrusive yet effective detection of radio-layer performance degradation, revealing the sensitivity of tail latency to UE type, distance, and network load. This facilitates practical and efficient fault localization and monitoring in O-RAN deployments.

cross-layer performanceO-RAN diagnosticsradio-layer dynamics

Current evaluations of large models predominantly rely on end-to-end metrics, which obscure the underlying causes of performance variations due to hardware and software configurations. This work proposes the first reproducible, execution-trace-based benchmarking framework that constructs a community-extensible, trace-level evidence ecosystem through fine-grained execution traces, YAML-based workload specifications, and containerized launch scripts. The framework enables in-depth analysis of computational, memory, and communication efficiency. Using this approach, the study systematically quantifies—for the first time—the impact of parallelization strategies, interconnect bandwidth, and framework-level optimizations on training performance. Key findings include: high compute-communication overlap does not necessarily reduce step time; doubling TPU interconnect bandwidth yields significantly greater benefits than on GPUs for small-to-medium workloads; and performance gaps of up to 3× exist between optimal configurations across different frameworks.

benchmarkingconfiguration spaceLLM infrastructure

Hot Scholars

IS

Ion Stoica

Professor of Computer Science, UC Berkeley
Cloud ComputingNetworkingDistributed SystemsBig Data
PA

Pablo Ameigeiras

Universidad de Granada
Wireless Communications and Networks
DK

Dragi Kimovski

Klagenfurt University
Parallel ProcessingCloud ComputingEdge computing