Score
Designs, implements, and analyzes instrumentation, tooling, and workflows to collect and interpret runtime performance metrics, traces, and hardware counters across processes, hosts, and distributed components; produces profiles and end-to-end or distributed traces to locate bottlenecks, performance regressions, and correctness issues. Builds and integrates profiling toolchains for workload and system-level characterization, debugging, tuning, and hardware-level analysis, and includes per-request or per-user–level profiling when needed to attribute observed behavior.
Modern software systems frequently exhibit application-level resource contention bottlenecks—such as blocking on custom application events—that evade detection by conventional performance profilers due to complex dependencies and bespoke resource management. To address this, we propose OmniResource Profiling, the first method to jointly leverage system-level metrics and application-level event-waiting relationships. It employs a lightweight LLM-assisted static analysis to automatically identify custom resources and cross-execution-trace runtime variable comparison for precise root-cause localization. Evaluated on 12 known performance issues across five real-world applications, OmniResource achieves 100% diagnostic accuracy and uncovers two previously undetected bottlenecks. Crucially, it requires no intrusive instrumentation, balancing high precision with practical deployability. This work delivers the first end-to-end solution for application-level resource contention analysis.
Resource management across the IoT/Edge/Cloud continuum faces a fundamental trade-off among prediction accuracy, real-time responsiveness, and deployment lightweightness; existing approaches rely either on runtime sampling or static rules, failing to reconcile these requirements. This paper proposes a metadata-driven runtime workload profiling framework. First, it formally defines the mapping between static metadata and dynamic runtime behavior. Second, it introduces a clustering-based mechanism for extracting semantically significant metadata features—without requiring execution traces or online profiling. Third, it integrates lightweight feature engineering with regression modeling to achieve low-overhead, high-accuracy resource demand prediction. Evaluated across Alibaba’s ML workloads and Google’s cluster traces, the framework maintains high prediction accuracy even under partial data anonymization, substantially outperforming conventional static prediction and online profiling baselines. It achieves superior real-time performance, prediction fidelity, and practical deployability.
Non-author engineers often struggle to associate performance bottlenecks with program semantics. Method: This paper proposes an interpretable optimization approach that jointly leverages runtime performance data and code semantics. It introduces CodeBERT—the first pre-trained code model—into performance profiling: fine-tuning it to generate fine-grained code summaries and aligning these with call-path-level performance profiles collected by Async Profiler for Java applications; hot paths and their semantic summaries are then co-visualized in a graphical interface. Contributions/Results: (1) We present the first semantic-augmented performance profiling framework built upon a pre-trained code model; (2) the approach significantly improves bottleneck interpretability and optimization guidance. Experiments across multiple Java benchmarks demonstrate that our system effectively reduces developers’ cognitive load, shortening average bottleneck localization time by 37.2%.
Programming heterogeneous exascale HPC systems is hindered by the proliferation of complex, incompatible programming models (e.g., CUDA, SYCL, OpenMP) and the lack of traceability across CPU/GPU execution contexts. Method: We propose the first semantic-aware, full-stack API tracing framework built upon LTTng kernel tracing, integrating user-space dynamic instrumentation with multi-model API signature parsing to capture fine-grained, low-overhead, configurable API call chains across hardware and programming abstractions. Contribution/Results: Unlike conventional tracers logging only function names and timestamps, our framework enables cross-vendor, cross-abstraction behavioral correlation and end-to-end call-chain reconstruction in real HPC applications. It accurately identifies cross-model performance bottlenecks and implementation flaws, improving debugging efficiency by over 3× and significantly enhancing portability and debuggability of heterogeneous programming models.
This work addresses the challenges of fragmented and heterogeneous performance monitoring units in modern heterogeneous multi-core SoCs, which complicate data collection, synchronization, and correlation. The paper proposes a centralized performance monitoring architecture—demonstrated for the first time in a RISC-V SoC—that enables unified cross-component monitoring. Hardware microarchitectural events are collected by Event Monitoring Units (EVUs) and routed via the AXI4 bus to an Advanced Performance Monitoring Unit (APMU) for consolidated processing. By integrating programmable counters and a dedicated processing unit, the architecture simplifies interface design, enhances event correlation capabilities, and supports event-driven software mechanisms. Experimental results validate its effectiveness in real-time resource management, application profiling, and counter attribution scenarios.
Existing MPI performance analysis tools rely on time-aggregated metrics, which obscure transient bottlenecks. To address this, we propose a fine-grained, time-windowed trace analysis method that partitions execution traces into fixed or adaptive temporal windows and computes time-resolved metrics—including communication efficiency, load balance, and serialization overhead. Our approach integrates Paraver-based post-processing, critical path reconstruction, and event anomaly correction (e.g., clock skew compensation and unmatched MPI event reconciliation) to enable high-precision localization of transient bottlenecks. Evaluation on real-world applications (LaMEM, ls1-MarDyn) and synthetic benchmarks demonstrates that our method significantly improves both accuracy and scalability in identifying transient performance issues within large-scale traces, thereby overcoming the inherent limitations of global aggregation-based analysis.
This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.
This work addresses the challenge of reliably detecting and interpreting anomalous node behaviors in large-scale high-performance computing (HPC) systems, where high-dimensional, unlabeled monitoring data complicates analysis. To tackle this, the authors propose a scalable, interactive visual analytics system that uniquely integrates contrastive learning with multi-resolution dynamic mode decomposition. Coupled with a two-stage dimensionality reduction pipeline and tailored visual encodings, the system enables users to explore, compare, and iteratively validate hypotheses about node behaviors. The approach effectively uncovers subtle intra- and inter-cluster behavioral differences, automatically identifying semantically meaningful node clusters in two real-world case studies. Expert evaluations demonstrate that the system substantially enhances both the accuracy and interpretability of anomaly detection in complex HPC environments.
This work addresses the limited feature coverage in high-performance computing caused by hardware performance counters constrained by the number of simultaneously collectible metrics. To overcome this limitation without relying on hardware multiplexing, the authors propose a heuristic multi-run execution trace merging method that aligns and fuses counter data collected across multiple program executions. By analyzing MPI communication structures, timing patterns, and behavioral characteristics, the approach constructs a high-dimensional, unified synthetic trace that expands the effective feature space. This enriched representation enables the training of more comprehensive machine learning–based performance models. Experimental evaluation on the MareNostrum5 platform demonstrates that the merged counters retain high accuracy and significantly improve performance prediction for diverse kernel functions and real-world applications.