low latency

Designs, builds, and evaluates systems, network components, algorithms, or hardware to minimize end-to-end latency and its variation (jitter), including measurement, profiling, and optimization of processing, buffering, queuing, scheduling, and transmission paths. Works to meet specified latency targets for real-time or interactive workloads by making trade-offs among throughput, resource usage, and consistency.

lowlatency

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.83
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$193K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses energy optimization for latency-sensitive network services under tail-latency SLA constraints. We propose a black-box, online co-tuning method that jointly optimizes interrupt coalescing (packet batching) and Dynamic Voltage and Frequency Scaling (DVFS), requiring no modifications to applications or the OS kernel. Our approach employs a generic, application- and system-agnostic Bayesian optimization controller that rapidly converges to the optimal energy-efficiency operating point under SLA constraints using minimal online probing. Key contributions include: (i) achieving up to 60% energy reduction while meeting tail-latency SLAs; (ii) revealing that specialized OS kernels improve energy efficiency by over 2× compared to general-purpose kernels; and (iii) demonstrating strong generalizability and stability across diverse hardware platforms. The method enables practical, deployment-ready energy-aware tuning for high-performance network services without sacrificing latency guarantees.

Achieve SLA targets efficientlyControl batching and processing rateOptimize energy-performance trade-offs

Characterization of latency and jitter in TSN emulation

Jun 02, 2025
AG
Alex Gracia
🏛️ Universidad de Zaragoza | Intel Corporation | CINVESTAV

Existing software simulation of Time-Sensitive Networking (TSN) suffers from insufficient accuracy in measuring bridge delay and jitter, undermining the fidelity and reproducibility of TSN emulation. Method: This paper introduces the first systematic timestamping methodology for TSN simulation on Linux/Mininet, rigorously evaluating four timestamping mechanisms—including SO_TIMESTAMPING—under TSN traffic shaped by Credit-Based Shaping (CBS) and Asynchronous Traffic Shaping (ATS). Leveraging configurable Mininet topologies, the approach integrates scheduling solution generation, deployment validation, and cross-platform optimization—supporting both Intel Time-Coordinated Computing (TCC)-enabled and -disabled modes on industrial PCs and workstations. Contribution/Results: The framework achieves sub-microsecond bridge delay characterization and, for the first time, experimentally validates end-to-end deterministic guarantees on real hardware. It overcomes critical bottlenecks in clock synchronization precision and scheduling fidelity, significantly enhancing the trustworthiness and reproducibility of TSN simulation.

Characterizing latency and jitter in TSN emulation environmentsEvaluating timestamping methods for TSN network traffic profilingSolving TSN scheduling challenges in software-based emulation

Characterizing the Impact of Active Queue Management on Speed Test Measurements

Nov 24, 2025
SR
Siddhant Ray
🏛️ University of Chicago | Cal Poly | ENS Lyon

Existing speed measurement tools focus on peak throughput and poorly reflect users’ perceived responsiveness; emerging metrics such as “latency under load” show promise but their sensitivity to Active Queue Management (AQM) configurations remains unclear. Method: We empirically evaluate three mainstream AQM schemes—CoDel, FQ-CoDel, and SFQ—in a controlled network environment, systematically analyzing their impact on throughput and latency distributions, particularly latency under load. Results: AQM significantly alters speed test outcomes, with distinct latency-throughput trade-offs observed across algorithms under high load. Current measurement platforms, if uncalibrated for AQM, yield misleading latency estimates, undermining the reliability of policy and regulatory decisions. This study is the first to quantitatively characterize the structural impact of AQM on emerging speed metrics, providing critical empirical evidence to inform standardization of measurement tools and evidence-based network governance.

Calibrating speed tests for accurate policy guidanceComparing throughput variance across different AQM schemesUnderstanding AQM's impact on speed test latency metrics

In 5G–Time-Sensitive Networking (TSN) convergence scenarios for Industry 4.0, stringent deterministic communication requirements are undermined by 5G channel jitter, which degrades the periodic scheduling performance of the IEEE 802.1Qbv Time-Aware Shaper (TAS). Method: This work quantifies, for the first time, the impact of 5G jitter on TAS scheduling within a real-world 5G–TSN heterogeneous testbed, and proposes a jitter mitigation strategy based on cross-layer alignment between TAS configuration parameters (e.g., gate control list cycle length) and 5G transmission characteristics (e.g., TTI duration and frame structure). Contribution/Results: Experimental evaluation demonstrates that optimized cycle alignment reduces end-to-end jitter to within ±10 μs—significantly outperforming misaligned configurations—and reliably satisfies deterministic latency requirements of industrial applications such as robotic control. This study provides a reproducible empirical foundation and an engineering pathway for cross-layer coordinated scheduling in 5G–TSN integration.

Analyzing 5G jitter's effect on TSN schedulingEvaluating 5G-TSN integration impact on determinismMitigating jitter to maintain deterministic performance

Latest Papers

What's happening recently
View more

In the era of sub-millisecond networking, host-side latencies—such as those introduced by the kernel network stack and application scheduling—have become the dominant bottleneck for end-to-end low-latency performance, yet production environments lack effective means for continuous monitoring. This work proposes and implements netstacklat, the first system to enable low-overhead, continuous end-to-end latency monitoring within the Linux kernel network stack. By leveraging lightweight kernel probes and an efficient performance monitoring framework, netstacklat accurately captures the data path latency from the network interface card to the application across 144 diverse Nginx/Apache HTTP workloads, incurring less than 6% overhead even at tail latencies. The tool has been successfully deployed across Cloudflare’s global CDN infrastructure, demonstrating its scalability and practical utility in real-world production settings.

host latencylatency monitoringnetwork stack

This work addresses the challenge of meeting latency constraints in multi-model large language model (LLM) serving, where the tight coupling between request routing and resource allocation renders traditional approaches ineffective. To tackle this issue, the paper presents the first joint optimization framework that simultaneously models both decisions. The approach constructs a deployment-aware latency model based on empirical system measurements and leverages a dual pricing mechanism to solve the constrained optimization problem under latency service-level objectives (SLOs). Experimental results demonstrate that, on the same GPU cluster, varying resource allocations can lead to up to an 87% difference in output quality, highlighting the critical importance of co-optimizing routing and resource provisioning. The proposed framework effectively enhances service quality while rigorously satisfying latency requirements.

GPU clusterslatency SLOmulti-model LLM serving

This work addresses the significant degradation in Quality of Service (QoS) in multi-robot systems caused by static task offloading under network latency, jitter, and edge resource contention. To mitigate these issues, the authors propose and implement an edge-based closed-loop Adaptive Task Placement (ATP) controller that dynamically orchestrates the placement of perception tasks between local robots and edge servers. ATP employs a lightweight, multi-metric-driven mechanism, leveraging a normalized multidimensional cost function over a two-second control window to jointly evaluate communication latency, CPU utilization, and migration overhead. Implemented on a real multi-robot platform, ATP achieves QoS-aware dynamic task orchestration for the first time in such settings. Experimental results demonstrate that under computational stress and network failure scenarios, ATP substantially reduces tail latency and deadline violation rates compared to static offloading strategies, offering practical design guidelines for cloud-edge robotic deployments.

Edge computingLatencyMulti-robot systems

This work addresses a critical limitation in existing latency models for strict-priority scheduling, which neglect the impact of the transmit ring buffer (TXR) at switch egress ports, leading to inaccurate delay estimates for high-priority packets. For the first time, this study incorporates TXR into the end-to-end delay analysis framework by proposing a method to measure TXR size and integrating deterministic and stochastic modeling, network measurements, and hardware behavior characterization. The resulting model significantly improves the accuracy of both worst-case delay bounds and delay distribution predictions for high-priority traffic. Validated across multiple commercial off-the-shelf switches, the approach effectively bridges the gap between theoretical models and real hardware behavior, thereby enhancing the reliability of real-time network system design.

Latency ModelingPacket DelayStrict Priority