Score
Designs and implements measurement, profiling, and benchmarking artifacts (testbeds, probes, trace collection, benchmark suites) to measure and characterize latency and throughput properties of systems and services in real-time or near-real-time settings, including end-to-end, network, and tail-latency behavior. Builds analytical and statistical models and tooling to estimate and predict latency distributions, performs latency and memory/throughput profiling, and conducts latency‑accuracy and other tradeoff analyses to validate real-time constraints and guide latency-focused optimization.
This work addresses the limitations of existing performance evaluation approaches for distributed computing continua, which often focus on a single dimension and fail to holistically characterize the behavior of cross-layer heterogeneous systems. The paper presents the first systematic framework that establishes a comprehensive taxonomy of performance metrics spanning three layers—computation, networking, and application/user—as well as emerging non-functional attributes such as sustainability and observability. By integrating mathematical modeling with cross-layer analysis, the study rigorously defines the applicability, measurement phases, and specifications for each metric category. The resulting framework is both clearly structured and extensible, offering a solid theoretical foundation and practical guidance for unified performance assessment in dynamic, heterogeneous environments.
Existing speed measurement tools focus on peak throughput and poorly reflect users’ perceived responsiveness; emerging metrics such as “latency under load” show promise but their sensitivity to Active Queue Management (AQM) configurations remains unclear. Method: We empirically evaluate three mainstream AQM schemes—CoDel, FQ-CoDel, and SFQ—in a controlled network environment, systematically analyzing their impact on throughput and latency distributions, particularly latency under load. Results: AQM significantly alters speed test outcomes, with distinct latency-throughput trade-offs observed across algorithms under high load. Current measurement platforms, if uncalibrated for AQM, yield misleading latency estimates, undermining the reliability of policy and regulatory decisions. This study is the first to quantitatively characterize the structural impact of AQM on emerging speed metrics, providing critical empirical evidence to inform standardization of measurement tools and evidence-based network governance.
In the era of sub-millisecond networking, host-side latencies—such as those introduced by the kernel network stack and application scheduling—have become the dominant bottleneck for end-to-end low-latency performance, yet production environments lack effective means for continuous monitoring. This work proposes and implements netstacklat, the first system to enable low-overhead, continuous end-to-end latency monitoring within the Linux kernel network stack. By leveraging lightweight kernel probes and an efficient performance monitoring framework, netstacklat accurately captures the data path latency from the network interface card to the application across 144 diverse Nginx/Apache HTTP workloads, incurring less than 6% overhead even at tail latencies. The tool has been successfully deployed across Cloudflare’s global CDN infrastructure, demonstrating its scalability and practical utility in real-world production settings.
Existing MPI performance analysis tools rely on time-aggregated metrics, which obscure transient bottlenecks. To address this, we propose a fine-grained, time-windowed trace analysis method that partitions execution traces into fixed or adaptive temporal windows and computes time-resolved metrics—including communication efficiency, load balance, and serialization overhead. Our approach integrates Paraver-based post-processing, critical path reconstruction, and event anomaly correction (e.g., clock skew compensation and unmatched MPI event reconciliation) to enable high-precision localization of transient bottlenecks. Evaluation on real-world applications (LaMEM, ls1-MarDyn) and synthetic benchmarks demonstrates that our method significantly improves both accuracy and scalability in identifying transient performance issues within large-scale traces, thereby overcoming the inherent limitations of global aggregation-based analysis.
Existing software simulation of Time-Sensitive Networking (TSN) suffers from insufficient accuracy in measuring bridge delay and jitter, undermining the fidelity and reproducibility of TSN emulation. Method: This paper introduces the first systematic timestamping methodology for TSN simulation on Linux/Mininet, rigorously evaluating four timestamping mechanisms—including SO_TIMESTAMPING—under TSN traffic shaped by Credit-Based Shaping (CBS) and Asynchronous Traffic Shaping (ATS). Leveraging configurable Mininet topologies, the approach integrates scheduling solution generation, deployment validation, and cross-platform optimization—supporting both Intel Time-Coordinated Computing (TCC)-enabled and -disabled modes on industrial PCs and workstations. Contribution/Results: The framework achieves sub-microsecond bridge delay characterization and, for the first time, experimentally validates end-to-end deterministic guarantees on real hardware. It overcomes critical bottlenecks in clock synchronization precision and scheduling fidelity, significantly enhancing the trustworthiness and reproducibility of TSN simulation.
This paper addresses the challenge of passively measuring end-to-end response latency under transport-layer encryption (e.g., TLS/QUIC), where application-layer headers and client instrumentation are unavailable. We propose PIRATE, a passive latency estimation algorithm that relies solely on observable client→server traffic. Its core innovation is the first use of causally linked request-pair time differences—without client-side instrumentation or plaintext headers—to accurately proxy application-layer round-trip latency, inherently supporting asymmetric routing. The method integrates causal request-pair detection, passive temporal modeling, and DSR-aware load-balancing coordination. Evaluated on real-world web services, PIRATE achieves ≤1% estimation error for client-side latency. When deployed at Layer-4 load balancers, it reduces tail latency by 37%, significantly enhancing bottleneck identification, adaptive scheduling, and attack mitigation capabilities.
This work addresses the lack of existing tools capable of continuous performance validation and regression detection across entire datacenter clusters. The authors propose the first cluster-wide continuous benchmarking framework that supports unified scheduling, enabling simultaneous task distribution to all nodes and systematic collection of multidimensional performance metrics—spanning CPU, GPU, memory, interconnects, I/O, power consumption, frequency, and temperature—across both space and time. This framework facilitates performance regression detection under software and hardware changes as well as analysis of hardware variability. Experiments on the NHR@FAU cluster reveal intra-node performance variations below 1% among identically configured nodes, while inter-node differences reach up to 5%. The study further uncovers, for the first time, significant disparities in the performance–power relationship between air-cooled and liquid-cooled nodes.
This study addresses the limited sensitivity of traditional cloud service performance regression detection, which is often hindered by I/O fluctuations and infrastructure changes. The authors propose a novel paradigm termed “Duet Instrumentation,” which uniquely integrates large language model (LLM)-driven code change analysis with synchronized dual-version benchmarking. By leveraging an LLM to precisely identify performance-relevant changes between consecutive versions, the method dynamically instruments only those critical code regions, achieving high-sensitivity regression detection with low overhead. Evaluated in real-world environments, the approach attains a precision of 58%, recall of 93%, and specificity of 71%, effectively detecting performance regressions as subtle as one-fifth the severity detectable by conventional methods.
This work addresses behavioral inconsistencies between Linux kernel implementations and userspace simulators of the L4S (Low Latency, Low Loss, Scalable throughput) mechanism, which hinder experimental reproducibility and parameter portability. We present the first scalable implementation of the DualPI2 active queue management algorithm in Mahimahi and conduct a systematic comparison against its kernel counterpart across diverse traffic patterns and network conditions. Through comprehensive behavioral characterization and parameter sensitivity analysis, we identify the bandwidth-delay product (BDP) as a critical factor governing cross-platform discrepancies. Our findings reveal specific parameter configurations that improve alignment under low-BDP scenarios, while also exposing persistent structural deviations under high load. This study provides both a practical simulation tool and empirical guidance for accurate L4S experimentation and deployment.
This work proposes a cross-layer, interpretable performance diagnosis method to address the challenge of detecting subtle radio-layer dynamic anomalies in O-RAN systems when end-to-end latency appears stable. Leveraging real-world measurements across multiple distances and user equipment (UE) types, the approach jointly analyzes application-layer tail latency—such as the 95th percentile—with radio-layer metrics including scheduling behavior, modulation and coding scheme (MCS), block error rate (BLER), and signal quality to construct lightweight “degradation flags.” The method enables non-intrusive yet effective detection of radio-layer performance degradation, revealing the sensitivity of tail latency to UE type, distance, and network load. This facilitates practical and efficient fault localization and monitoring in O-RAN deployments.
Current evaluations of large models predominantly rely on end-to-end metrics, which obscure the underlying causes of performance variations due to hardware and software configurations. This work proposes the first reproducible, execution-trace-based benchmarking framework that constructs a community-extensible, trace-level evidence ecosystem through fine-grained execution traces, YAML-based workload specifications, and containerized launch scripts. The framework enables in-depth analysis of computational, memory, and communication efficiency. Using this approach, the study systematically quantifies—for the first time—the impact of parallelization strategies, interconnect bandwidth, and framework-level optimizations on training performance. Key findings include: high compute-communication overlap does not necessarily reduce step time; doubling TPU interconnect bandwidth yields significantly greater benefits than on GPUs for small-to-medium workloads; and performance gaps of up to 3× exist between optimal configurations across different frameworks.