Score
Designs and implements benchmarks and evaluation pipelines that measure and compare trade-offs between system efficiency (latency, throughput, compute, memory, energy, cost) and output quality (accuracy, fidelity, robustness, utility). Builds measurement protocols, metrics, data sampling and experiments to produce comparable efficiency–quality curves and Pareto analyses, and analyzes results to guide system configuration, optimization, or component selection.
The high-performance computing (HPC) domain suffers from an abundance of benchmarking tools and the absence of a standardized, unified classification framework. Method: This paper proposes the first standardized benchmark taxonomy for HPC, derived from a systematic literature review and multi-dimensional feature analysis across hardware, software, and algorithmic layers. A structured classification model is constructed, with key attributes—including target workload, portability, scalability, and measurement granularity—concisely tabulated. An interactive web-based platform is further developed to enable dynamic, dimension-driven querying, cross-benchmark comparison, and visual analytics. Contribution/Results: The taxonomy systematically organizes over 100 mainstream HPC benchmarks, significantly enhancing efficiency and consistency for architects, researchers, and scientific users in system evaluation, benchmark selection, and performance optimization. It establishes a foundational framework for standardizing HPC performance assessment and facilitates reproducible, comparable, and interpretable benchmarking practices.
AI infrastructure confronts multidimensional physical and economic constraints—including power, thermal management, water usage, interconnect bandwidth, memory capacity, and data throughput—while existing metrics (e.g., PUE, TCO) are siloed and fail to capture the coupled trade-offs among energy efficiency, performance, and cost, hindering cross-layer co-optimization. To address this, we propose a unified measurement architecture grounded in a 6×3 cross-layer taxonomy—spanning facility, network, compute, storage, software, and application layers, each annotated with physical, computational, and economic semantics—and introduce the Measurement Propagation Graph (MPG) to enable, for the first time, system-level, three-dimensional relational modeling. Leveraging systematic literature review, meta-analysis, and graph-based modeling, our framework integrates heterogeneous, multi-source metrics. It supports benchmarking, capacity planning, and total cost of ownership analysis, substantially enhancing interpretability of AI cluster efficiency frontiers and enabling rigorous multi-objective optimization.
This work addresses the limitations of existing performance evaluation approaches for distributed computing continua, which often focus on a single dimension and fail to holistically characterize the behavior of cross-layer heterogeneous systems. The paper presents the first systematic framework that establishes a comprehensive taxonomy of performance metrics spanning three layers—computation, networking, and application/user—as well as emerging non-functional attributes such as sustainability and observability. By integrating mathematical modeling with cross-layer analysis, the study rigorously defines the applicability, measurement phases, and specifications for each metric category. The resulting framework is both clearly structured and extensible, offering a solid theoretical foundation and practical guidance for unified performance assessment in dynamic, heterogeneous environments.
Existing CPU benchmarks (e.g., SPEC CPU2017) lack explicit system configuration specifications, leading to performance interference from non-CPU components and severely undermining comparability, consistency, and reproducibility. Method: We propose a novel CPU performance evaluation paradigm grounded in the principle of “fully specified and valid configurations,” establishing a systematic modeling framework that spans the complete configuration space; we design an unbiased sampling strategy that uniformly weights all compliant configurations; and we replace point estimates with confidence intervals and associated confidence levels for performance reporting. Results: Experiments reveal up to 74.49× performance variation for the same CPU across compliant configurations. Our framework eliminates configuration ambiguity entirely, enabling fair cross-CPU comparisons and significantly improving consistency, reproducibility, and statistical rigor of benchmarking outcomes.
本文通过引入一个多维度度量框架,解决多尺度高性能计算中的资源管理和可持续性问题,指导现代工作负载的部署策略。
Traditional efficiency metrics struggle to accurately assess resource utilization in heterogeneous high-performance computing systems that combine CPUs and accelerators. This work extends the POP efficiency model by introducing a hardware-agnostic, host-device dual-branch hierarchical efficiency framework. It uniquely defines a multiplicative efficiency decomposition on the device side, symmetric to that on the host, separately capturing mixed execution/offload efficiency and device parallel efficiency. Implemented via the lightweight TALP monitoring library, the approach supports both runtime and post-mortem analysis and outputs results in human-readable and machine-readable formats. Experiments on synthetic benchmarks and three real-world HPC applications demonstrate that the proposed methodology effectively uncovers performance bottlenecks related to offloading, load balancing, and task scheduling, offering developers actionable insights for optimization.
Cloud-native applications operate in multi-tenant, shared cloud environments, where conventional energy-efficiency evaluation methods—relying solely on isolated local metrics (e.g., CPU utilization)—fail to capture holistic system-level energy behavior. To address this, we propose the first automated, scalable energy-efficiency experimentation framework tailored for Kubernetes-native applications, enabling joint quantification of energy consumption and QoS metrics (latency, throughput, error rate) across container, platform, and infrastructure layers. The framework tightly integrates eBPF for fine-grained observability, Prometheus for metric collection, and hardware power interfaces (RAPL/ACPI) for accurate energy measurement, and establishes an end-to-end sustainability assessment pipeline. Evaluation across multiple open-source cloud-native applications demonstrates up to 42% energy-efficiency variation among architectural variants, while precisely characterizing their trade-offs with P99 latency and service availability.
This study addresses the high variability and low reliability of performance benchmarking results for stream-processing applications in cloud environments. Over three months, we conducted a large-scale longitudinal empirical study across multiple geographic regions and heterogeneous hardware—including diverse CPU architectures. Leveraging Kubernetes-based automated deployment, high-frequency repeated benchmarking, and time-series statistical analysis, we systematically characterized end-to-end cloud performance variability for the first time at the application level. We discovered that variability exhibits statistically significant diurnal and weekly periodicity (amplitude ≤2.5%) and a coefficient of variation <3.7%—substantially lower than commonly assumed in industry. Moreover, infrastructure sharing incurs at most a 2.5-percentage-point loss in measurement precision. These findings demonstrate strong robustness across regions and CPU architectures, providing empirical evidence and methodological foundations for enhancing reproducibility and trustworthiness in cloud-native benchmarking.
本文通过对比SPEC CPU2026基准测试与其上游开源版本在单副本和多副本运行场景下的性能,量化分析了两者之间的'保真度差距',验证了SPEC方法的有效性。
This study addresses the longstanding fragmentation in microservice energy efficiency research, which has been siloed across runtime, infrastructure, and architectural layers, lacking a unified lifecycle perspective and consistent measurement methodology. Employing Kitchenham’s systematic literature review approach—augmented by searches across four major databases and snowballing techniques—the authors analyze 40 core studies to integrate multidimensional viewpoints for the first time. Their synthesis reveals an overwhelming emphasis on runtime optimizations, such as scheduling and resource management, typically relying on coarse-grained monitoring and model-based estimations, while largely neglecting energy-aware integration during architectural design and fine-grained measurement practices. The work establishes energy efficiency as a critical cross-cutting architectural attribute throughout the microservice lifecycle and underscores the urgent need for early-design support and a unified measurement framework.
This study addresses the challenge of optimizing server energy efficiency in high-throughput computing environments, where performance and energy consumption are often at odds. Leveraging real-world operational data and targeted experiments, the work systematically investigates how server configurations influence power consumption, performance, and carbon emissions, uncovering key barriers to implementing effective energy-saving measures in practice. Through empirical power monitoring, workload modeling, and carbon footprint assessment, the authors identify critical factors governing energy efficiency and propose a practical configuration strategy that simultaneously ensures performance guarantees and advances low-carbon objectives. Evaluated under representative high-throughput workloads, the proposed approach achieves substantial reductions in both energy use and carbon emissions.
This study addresses the absence of open standards for CPU pipeline visualization tools and the difficulty in localizing performance bottlenecks. To this end, it proposes an open-source event stream format alongside Catscan, an interactive viewer. Methodologically, this work introduces a structured event stream based on transactional relationships, integrating typed event modeling, persistent highlighting techniques, and domain-specific search algorithms to enable microarchitectural trace analysis from symptoms down to individual instructions. Furthermore, it supports resource-oriented views synchronized with comparative trace alignment. By successfully reproducing industry-grade debugging workflows, this project provides the community with production-validated microarchitectural visualization infrastructure.
研究通过构建成本、质量和延迟帕累托图谱,评估54种Qwen2.5-7B-Instruct配置在不同GPU上的表现,以找到最佳部署方案。