Score
Designs, implements, and evaluates low-latency systems and architectures, including networking and serving stacks, along with the profiling, monitoring, benchmarking, and tooling needed to measure and enforce latency budgets and cost–latency tradeoffs. Analyzes traces and profiles to find latency hotspots and applies optimizations across software, hardware, I/O, concurrency, serialization, caching, batching, and deployment/configuration to reduce average and tail latencies.
This work addresses the limitations of existing performance evaluation approaches for distributed computing continua, which often focus on a single dimension and fail to holistically characterize the behavior of cross-layer heterogeneous systems. The paper presents the first systematic framework that establishes a comprehensive taxonomy of performance metrics spanning three layers—computation, networking, and application/user—as well as emerging non-functional attributes such as sustainability and observability. By integrating mathematical modeling with cross-layer analysis, the study rigorously defines the applicability, measurement phases, and specifications for each metric category. The resulting framework is both clearly structured and extensible, offering a solid theoretical foundation and practical guidance for unified performance assessment in dynamic, heterogeneous environments.
This work addresses energy optimization for latency-sensitive network services under tail-latency SLA constraints. We propose a black-box, online co-tuning method that jointly optimizes interrupt coalescing (packet batching) and Dynamic Voltage and Frequency Scaling (DVFS), requiring no modifications to applications or the OS kernel. Our approach employs a generic, application- and system-agnostic Bayesian optimization controller that rapidly converges to the optimal energy-efficiency operating point under SLA constraints using minimal online probing. Key contributions include: (i) achieving up to 60% energy reduction while meeting tail-latency SLAs; (ii) revealing that specialized OS kernels improve energy efficiency by over 2× compared to general-purpose kernels; and (iii) demonstrating strong generalizability and stability across diverse hardware platforms. The method enables practical, deployment-ready energy-aware tuning for high-performance network services without sacrificing latency guarantees.
In the era of sub-millisecond networking, host-side latencies—such as those introduced by the kernel network stack and application scheduling—have become the dominant bottleneck for end-to-end low-latency performance, yet production environments lack effective means for continuous monitoring. This work proposes and implements netstacklat, the first system to enable low-overhead, continuous end-to-end latency monitoring within the Linux kernel network stack. By leveraging lightweight kernel probes and an efficient performance monitoring framework, netstacklat accurately captures the data path latency from the network interface card to the application across 144 diverse Nginx/Apache HTTP workloads, incurring less than 6% overhead even at tail latencies. The tool has been successfully deployed across Cloudflare’s global CDN infrastructure, demonstrating its scalability and practical utility in real-world production settings.
Existing MPI performance analysis tools rely on time-aggregated metrics, which obscure transient bottlenecks. To address this, we propose a fine-grained, time-windowed trace analysis method that partitions execution traces into fixed or adaptive temporal windows and computes time-resolved metrics—including communication efficiency, load balance, and serialization overhead. Our approach integrates Paraver-based post-processing, critical path reconstruction, and event anomaly correction (e.g., clock skew compensation and unmatched MPI event reconciliation) to enable high-precision localization of transient bottlenecks. Evaluation on real-world applications (LaMEM, ls1-MarDyn) and synthetic benchmarks demonstrates that our method significantly improves both accuracy and scalability in identifying transient performance issues within large-scale traces, thereby overcoming the inherent limitations of global aggregation-based analysis.
To address low utilization of deep idle states in latency-sensitive applications, this paper identifies a significant gap between theoretically available idle opportunities and their actual exploitation—caused by inaccurate idle scheduling decisions and non-negligible deep-sleep transition latency. We propose a queueing-theoretic modeling framework that integrates M/M/1, c×M/M/1, and M/M/c models, calibrated with real-world server workload traces, to quantify system-level idle potential under diverse configurations. For the first time, we systematically identify numerous untriggered deep-idle entry opportunities and develop a scalable methodology for idle-efficiency evaluation. The framework provides quantifiable, early-stage guidance for hardware–OS co-design, enabling energy-efficiency optimization and supporting system-level power management strategies that explicitly balance latency constraints and energy savings.
To address inefficient data migration and inaccurate performance prediction in heterogeneous storage systems (NVMe cache + HDD backend), this paper designs and implements a distributed two-tier storage system. We propose an online reinforcement learning–based dynamic data tiering scheduling algorithm and develop an end-to-end performance model integrating queuing network theory with fine-grained device behavior modeling. Our key contribution is the first scalable, fine-grained device behavior modeling method tailored for heterogeneous storage—enabling adaptive tiering management and precise performance prediction under high-concurrency I/O workloads in multi-core clusters. Experimental evaluation on multi-node clusters demonstrates an average model prediction error of less than 8%, a 27% improvement in I/O throughput, and a 34% reduction in average access latency. The framework provides a reusable modeling and optimization foundation for two-tier storage systems.
This work challenges the conventional focus on core utilization in resource management, which often overlooks the practical performance constraints imposed by power and thermal limits in modern multicore processors. Instead, it proposes a new paradigm centered on power budgeting, elevating idle-core waiting strategies to first-class design considerations. Rather than aggressively reclaiming idle cores—a practice that frequently overestimates benefits and incurs substantial scheduling overhead—the approach leverages efficient waiting mechanisms to release redistributable compute capacity. Empirical analysis on AMD EPYC platforms, accounting for processor topology, idle duration, and waiting policies, demonstrates that such strategies achieve a superior trade-off between energy efficiency and performance, offering greater practical advantages in real-world systems.
This paper identifies a critical gap in high-performance data transfer research: an overemphasis on network bandwidth while neglecting end-to-end bottlenecks—including latency, TCP congestion control, host CPU limitations, and virtualization—leading to severe discrepancies between benchmark results and real-world production performance. To address this, the authors propose a hardware–software co-design paradigm and develop a latency-programmable testbed. Leveraging high-fidelity wide-area network (WAN) modeling and cross-continental 100 Gbps measurements (Switzerland–California), they systematically isolate key constraints at the network edge and host side. Results demonstrate that primary bottlenecks reside predominantly at the network edge—not the core—and that stable, predictable data movement is achieved across 1–100+ Gbps. This significantly enhances performance fidelity in complex, heterogeneous environments.
This work addresses the challenge of meeting latency constraints in multi-model large language model (LLM) serving, where the tight coupling between request routing and resource allocation renders traditional approaches ineffective. To tackle this issue, the paper presents the first joint optimization framework that simultaneously models both decisions. The approach constructs a deployment-aware latency model based on empirical system measurements and leverages a dual pricing mechanism to solve the constrained optimization problem under latency service-level objectives (SLOs). Experimental results demonstrate that, on the same GPU cluster, varying resource allocations can lead to up to an 87% difference in output quality, highlighting the critical importance of co-optimizing routing and resource provisioning. The proposed framework effectively enhances service quality while rigorously satisfying latency requirements.