performance profiling

Designs, implements, and analyzes instrumentation, tooling, and workflows to collect and interpret runtime performance metrics, traces, and hardware counters across processes, hosts, and distributed components; produces profiles and end-to-end or distributed traces to locate bottlenecks, performance regressions, and correctness issues. Builds and integrates profiling toolchains for workload and system-level characterization, debugging, tuning, and hardware-level analysis, and includes per-request or per-user–level profiling when needed to attribute observed behavior.

performanceprofiling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.45
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$209K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

gigiProfiler: Diagnosing Performance Issues by Uncovering Application Resource Bottlenecks

Jul 08, 2025
YH
Yigong Hu
🏛️ Boston University | University of Washington | University of California, Los Angeles

Modern software systems frequently exhibit application-level resource contention bottlenecks—such as blocking on custom application events—that evade detection by conventional performance profilers due to complex dependencies and bespoke resource management. To address this, we propose OmniResource Profiling, the first method to jointly leverage system-level metrics and application-level event-waiting relationships. It employs a lightweight LLM-assisted static analysis to automatically identify custom resources and cross-execution-trace runtime variable comparison for precise root-cause localization. Evaluated on 12 known performance issues across five real-world applications, OmniResource achieves 100% diagnostic accuracy and uncovers two previously undetected bottlenecks. Crucially, it requires no intrusive instrumentation, balancing high precision with practical deployability. This work delivers the first end-to-end solution for application-level resource contention analysis.

Diagnosing performance bottlenecks in complex applicationsIdentifying application-level resource contention issuesUncovering root causes of performance problems

Formal and Empirical Study of Metadata-Based Profiling for Resource Management in the Computing Continuum

Apr 29, 2025
AM
Andrea Morichetta
🏛️ Vienna University of Technology | Universitat Pompeu Fabra

Resource management across the IoT/Edge/Cloud continuum faces a fundamental trade-off among prediction accuracy, real-time responsiveness, and deployment lightweightness; existing approaches rely either on runtime sampling or static rules, failing to reconcile these requirements. This paper proposes a metadata-driven runtime workload profiling framework. First, it formally defines the mapping between static metadata and dynamic runtime behavior. Second, it introduces a clustering-based mechanism for extracting semantically significant metadata features—without requiring execution traces or online profiling. Third, it integrates lightweight feature engineering with regression modeling to achieve low-overhead, high-accuracy resource demand prediction. Evaluated across Alibaba’s ML workloads and Google’s cluster traces, the framework maintains high prediction accuracy even under partial data anonymization, substantially outperforming conventional static prediction and online profiling baselines. It achieves superior real-time performance, prediction fidelity, and practical deployability.

Combining past execution traces with metadata for fast workload associationOptimizing resource distribution in IoT-Edge-Cloud computing continuumProfiling workload using static metadata for resource allocation

Non-author engineers often struggle to associate performance bottlenecks with program semantics. Method: This paper proposes an interpretable optimization approach that jointly leverages runtime performance data and code semantics. It introduces CodeBERT—the first pre-trained code model—into performance profiling: fine-tuning it to generate fine-grained code summaries and aligning these with call-path-level performance profiles collected by Async Profiler for Java applications; hot paths and their semantic summaries are then co-visualized in a graphical interface. Contributions/Results: (1) We present the first semantic-augmented performance profiling framework built upon a pre-trained code model; (2) the approach significantly improves bottleneck interpretability and optimization guidance. Experiments across multiple Java benchmarks demonstrate that our system effectively reduces developers’ cognitive load, shortening average bottleneck localization time by 37.2%.

Enhancing profiler usability with deep learningInterpreting complex performance data from profilersLinking inefficiencies to program semantics automatically

THAPI: Tracing Heterogeneous APIs

Mar 22, 2025
SB
Solomon Bekele
🏛️ Argonne National Laboratory | University De Los Andes

Programming heterogeneous exascale HPC systems is hindered by the proliferation of complex, incompatible programming models (e.g., CUDA, SYCL, OpenMP) and the lack of traceability across CPU/GPU execution contexts. Method: We propose the first semantic-aware, full-stack API tracing framework built upon LTTng kernel tracing, integrating user-space dynamic instrumentation with multi-model API signature parsing to capture fine-grained, low-overhead, configurable API call chains across hardware and programming abstractions. Contribution/Results: Unlike conventional tracers logging only function names and timestamps, our framework enables cross-vendor, cross-abstraction behavioral correlation and end-to-end call-chain reconstruction in real HPC applications. It accurately identifies cross-model performance bottlenecks and implementation flaws, improving debugging efficiency by over 3× and significantly enhancing portability and debuggability of heterogeneous programming models.

Capturing detailed API calls across HPC software stack layersDebugging and optimizing performance in heterogeneous computing environmentsUnderstanding interactions between multiple programming models in HPC systems

Latest Papers

What's happening recently
View more

This work addresses the challenges of fragmented and heterogeneous performance monitoring units in modern heterogeneous multi-core SoCs, which complicate data collection, synchronization, and correlation. The paper proposes a centralized performance monitoring architecture—demonstrated for the first time in a RISC-V SoC—that enables unified cross-component monitoring. Hardware microarchitectural events are collected by Event Monitoring Units (EVUs) and routed via the AXI4 bus to an Advanced Performance Monitoring Unit (APMU) for consolidated processing. By integrating programmable counters and a dedicated processing unit, the architecture simplifies interface design, enhances event correlation capabilities, and supports event-driven software mechanisms. Experimental results validate its effectiveness in real-time resource management, application profiling, and counter attribution scenarios.

Embedded SystemsEvent CorrelationHardware Performance Counters

Trace-based, time-resolved analysis of MPI application performance using standard metrics

Dec 01, 2025
KH
Kingshuk Haldar
🏛️ High Performance Computing Center Stuttgart | University of Stuttgart

Existing MPI performance analysis tools rely on time-aggregated metrics, which obscure transient bottlenecks. To address this, we propose a fine-grained, time-windowed trace analysis method that partitions execution traces into fixed or adaptive temporal windows and computes time-resolved metrics—including communication efficiency, load balance, and serialization overhead. Our approach integrates Paraver-based post-processing, critical path reconstruction, and event anomaly correction (e.g., clock skew compensation and unmatched MPI event reconciliation) to enable high-precision localization of transient bottlenecks. Evaluation on real-world applications (LaMEM, ls1-MarDyn) and synthetic benchmarks demonstrates that our method significantly improves both accuracy and scalability in identifying transient performance issues within large-scale traces, thereby overcoming the inherent limitations of global aggregation-based analysis.

Analyzes MPI application performance using time-resolved standard metricsIdentifies transient bottlenecks hidden by time-aggregated metrics in toolsProcesses execution traces robustly despite anomalies and large sizes

This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.

Autonomous AgentsHigh Performance ComputingJob Specification Translation

This work addresses the challenge of reliably detecting and interpreting anomalous node behaviors in large-scale high-performance computing (HPC) systems, where high-dimensional, unlabeled monitoring data complicates analysis. To tackle this, the authors propose a scalable, interactive visual analytics system that uniquely integrates contrastive learning with multi-resolution dynamic mode decomposition. Coupled with a two-stage dimensionality reduction pipeline and tailored visual encodings, the system enables users to explore, compare, and iteratively validate hypotheses about node behaviors. The approach effectively uncovers subtle intra- and inter-cluster behavioral differences, automatically identifying semantically meaningful node clusters in two real-world case studies. Expert evaluations demonstrate that the system substantially enhances both the accuracy and interpretability of anomaly detection in complex HPC environments.

anomaly detectionhigh-dimensional datahigh-performance computing

This work addresses the limited feature coverage in high-performance computing caused by hardware performance counters constrained by the number of simultaneously collectible metrics. To overcome this limitation without relying on hardware multiplexing, the authors propose a heuristic multi-run execution trace merging method that aligns and fuses counter data collected across multiple program executions. By analyzing MPI communication structures, timing patterns, and behavioral characteristics, the approach constructs a high-dimensional, unified synthetic trace that expands the effective feature space. This enriched representation enables the training of more comprehensive machine learning–based performance models. Experimental evaluation on the MareNostrum5 platform demonstrates that the merged counters retain high accuracy and significantly improve performance prediction for diverse kernel functions and real-world applications.

execution tracesfeature coveragehardware counters

Hot Scholars

HB

Holger Boche

Technische Universität München
Information TheorySignal ProcessingCommunication Theory
MZ

Matei Zaharia

UC Berkeley and Databricks
Distributed SystemsMachine LearningDatabasesSecurity
SL

Siqiang Luo

Assistant Professor of Nanyang Technological University
databasegraph data managementkey-value stores
MX

Minxian Xu

Associate Professor, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences
Cloud ComputingMicroservicesLLM Inference