Score
Analyze hardware and system architectures together with software behavior to locate and quantify resources (CPU cores, vector units, clock frequency, memory bandwidth, caches, and I/O) that limit application throughput or latency and to attribute slowdowns to causes such as cache contention, bandwidth saturation, or lack of effective vectorization. Produce measurements and models (profiling, microbenchmarks, performance counters) and recommend concrete mitigations (code or data-layout changes, compiler options, or hardware/configuration adjustments) to remove or reduce the bottleneck.
The high-performance computing (HPC) domain suffers from an abundance of benchmarking tools and the absence of a standardized, unified classification framework. Method: This paper proposes the first standardized benchmark taxonomy for HPC, derived from a systematic literature review and multi-dimensional feature analysis across hardware, software, and algorithmic layers. A structured classification model is constructed, with key attributes—including target workload, portability, scalability, and measurement granularity—concisely tabulated. An interactive web-based platform is further developed to enable dynamic, dimension-driven querying, cross-benchmark comparison, and visual analytics. Contribution/Results: The taxonomy systematically organizes over 100 mainstream HPC benchmarks, significantly enhancing efficiency and consistency for architects, researchers, and scientific users in system evaluation, benchmark selection, and performance optimization. It establishes a foundational framework for standardizing HPC performance assessment and facilitates reproducible, comparable, and interpretable benchmarking practices.
Accurately identifying computational, memory bandwidth, and memory latency bottlenecks—and quantifying associated resource slack—is critical yet challenging for HPC application performance tuning. This paper introduces the first model-agnostic, instruction-level precise noise-injection framework for bottleneck analysis. Leveraging the LLVM toolchain, it selectively injects computational or memory-access noise instructions to decouple the impact of each resource constraint, enabling fine-grained bottleneck classification and quantitative slack measurement. Unlike prior approaches, it requires no hardware modeling assumptions and is portable across diverse architectures. Evaluated on heterogeneous memory systems—including HBM and DDR—it demonstrates robust effectiveness. The method significantly improves the precision of optimization decisions and hardware selection guidance, addressing key limitations of existing tools in slack quantification and root-cause attribution of performance bottlenecks.
This study addresses the absence of open standards for CPU pipeline visualization tools and the difficulty in localizing performance bottlenecks. To this end, it proposes an open-source event stream format alongside Catscan, an interactive viewer. Methodologically, this work introduces a structured event stream based on transactional relationships, integrating typed event modeling, persistent highlighting techniques, and domain-specific search algorithms to enable microarchitectural trace analysis from symptoms down to individual instructions. Furthermore, it supports resource-oriented views synchronized with comparative trace alignment. By successfully reproducing industry-grade debugging workflows, this project provides the community with production-validated microarchitectural visualization infrastructure.
Modern software systems frequently exhibit application-level resource contention bottlenecks—such as blocking on custom application events—that evade detection by conventional performance profilers due to complex dependencies and bespoke resource management. To address this, we propose OmniResource Profiling, the first method to jointly leverage system-level metrics and application-level event-waiting relationships. It employs a lightweight LLM-assisted static analysis to automatically identify custom resources and cross-execution-trace runtime variable comparison for precise root-cause localization. Evaluated on 12 known performance issues across five real-world applications, OmniResource achieves 100% diagnostic accuracy and uncovers two previously undetected bottlenecks. Crucially, it requires no intrusive instrumentation, balancing high precision with practical deployability. This work delivers the first end-to-end solution for application-level resource contention analysis.
本文通过分析SPEC CPU 2026在AMD EPYC 9755处理器上的性能特征,采用多角度方法识别系统瓶颈和工作负载行为差异。
To address low per-GPU utilization and suboptimal hardware return-on-investment in heterogeneous multi-GPU systems, this paper proposes a data-driven analytical framework that establishes, for the first time, interpretable correlations between optimization strategies and GPU resource usage patterns. Our method integrates hardware performance counter profiling, multi-objective correlation modeling, and scientific proxy application benchmarks to construct a multidimensional metric suite characterizing application-device interaction behaviors. Unlike prior work—which focuses primarily on performance gains—our approach systematically uncovers the underlying mechanisms by which optimizations affect resource occupancy and utilization. Experimental evaluation on proxy applications demonstrates a 29.6% reduction in execution time, a 5.3% increase in average GPU utilization, and a 26.5% decrease in power consumption. These results establish a novel paradigm for resource-efficient, co-optimized heterogeneous accelerator systems.
研究通过结合静态代码特征、动态执行轨迹和内核级资源数据,发现13种性能原型,并提出一个多信号回归检测框架,以提高软件性能分析和预测的准确性。
This work addresses performance degradation in online data-intensive applications, which often stems from workload fluctuations and resource contention but remains elusive to conventional thread-state analysis due to complex cross-thread dependencies. The authors propose an application-agnostic diagnostic approach that leverages eBPF to collect 16 fine-grained metrics across six kernel subsystems—scheduling, VFS, networking, futex, multiplexed I/O, and block device I/O—and integrates a selective thread-tracing algorithm to precisely trace from entry-point threads to bottlenecked resources. By jointly modeling thread dynamics and resource interaction patterns, the method uniquely captures the propagation pathways of performance degradation. It enables low-overhead diagnosis of CPU, disk, lock, and external service contention across diverse workloads while uncovering internal application bottlenecks.
This work challenges the conventional focus on core utilization in resource management, which often overlooks the practical performance constraints imposed by power and thermal limits in modern multicore processors. Instead, it proposes a new paradigm centered on power budgeting, elevating idle-core waiting strategies to first-class design considerations. Rather than aggressively reclaiming idle cores—a practice that frequently overestimates benefits and incurs substantial scheduling overhead—the approach leverages efficient waiting mechanisms to release redistributable compute capacity. Empirical analysis on AMD EPYC platforms, accounting for processor topology, idle duration, and waiting policies, demonstrates that such strategies achieve a superior trade-off between energy efficiency and performance, offering greater practical advantages in real-world systems.
This work addresses the challenges of fragmented and heterogeneous performance monitoring units in modern heterogeneous multi-core SoCs, which complicate data collection, synchronization, and correlation. The paper proposes a centralized performance monitoring architecture—demonstrated for the first time in a RISC-V SoC—that enables unified cross-component monitoring. Hardware microarchitectural events are collected by Event Monitoring Units (EVUs) and routed via the AXI4 bus to an Advanced Performance Monitoring Unit (APMU) for consolidated processing. By integrating programmable counters and a dedicated processing unit, the architecture simplifies interface design, enhances event correlation capabilities, and supports event-driven software mechanisms. Experimental results validate its effectiveness in real-time resource management, application profiling, and counter attribution scenarios.
Existing performance analysis tools struggle to simultaneously capture temporal dynamics and a holistic view of performance bottlenecks: Roofline models neglect time evolution, while profilers and tracers obscure theoretical performance limits. This work proposes campaign diagrams—a novel visualization framework that uniquely integrates temporal phases with multidimensional resource utilization, including computational throughput, memory bandwidth, data traffic, and latency. Campaign diagrams can be generated from analytical models, simulations, or profiling data, concurrently displaying both theoretical performance ceilings and achieved performance. The approach uncovers cross-phase optimization opportunities often missed by conventional tools, such as counterintuitive cases where enhancing low-intensity operators improves end-to-end performance. Validated on low-rank GEMM and Mamba workloads, the method successfully identifies potential for operator fusion and pipeline optimizations, demonstrating its efficacy in diagnosing deep-rooted performance bottlenecks.