profile and optimize performance

Designs and builds instrumentation, profiling and monitoring pipelines to measure runtime, memory, and hardware-related resource use, and produces normalized metrics and visualizations for comparative and statistical analysis. Develops analytical, statistical, and probabilistic performance models and predictors to perform scalability, trade-off, and optimization analyses, implements targeted code- and layout-level optimizations, and derives bounds or guarantees on performance.

profileandoptimizeperformance

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.13
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$218K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Opal: A Modular Framework for Optimizing Performance using Analytics and LLMs

Oct 01, 2025
MZ
Mohammad Zaeed
🏛️ Texas State University | University of Novi Sad

Large language models (LLMs) struggle to autonomously optimize GPU kernel code without performance context, while conventional profiling tools identify bottlenecks but fail to generate executable optimization strategies. Method: This paper introduces the first GPU kernel auto-optimization framework integrating LLMs with dynamic hardware performance insights—including hardware counter metrics and Roofline model analysis—via structured prompting that precisely encodes bottleneck characteristics to guide high-fidelity, executable code generation. Contribution/Results: Evaluated across 1,640 experiments, the framework achieves performance improvements in over 98.5% of cases, with average speedups ranging from 19.34% to 52.3%; generated code exhibits near-perfect functional correctness. This work bridges the longstanding gap between low-level performance analysis and high-level optimization decision-making, establishing a reproducible, scalable paradigm for AI-driven, system-level GPU code optimization.

Automating code optimization using LLMs with performance contextBridging performance analysis insights to actionable optimization decisionsGenerating trustworthy GPU kernel optimizations via analytics-guided LLMs

This work addresses the challenges of fragmented and heterogeneous performance monitoring units in modern heterogeneous multi-core SoCs, which complicate data collection, synchronization, and correlation. The paper proposes a centralized performance monitoring architecture—demonstrated for the first time in a RISC-V SoC—that enables unified cross-component monitoring. Hardware microarchitectural events are collected by Event Monitoring Units (EVUs) and routed via the AXI4 bus to an Advanced Performance Monitoring Unit (APMU) for consolidated processing. By integrating programmable counters and a dedicated processing unit, the architecture simplifies interface design, enhances event correlation capabilities, and supports event-driven software mechanisms. Experimental results validate its effectiveness in real-time resource management, application profiling, and counter attribution scenarios.

Embedded SystemsEvent CorrelationHardware Performance Counters

From Profiling to Optimization: Unveiling the Profile Guided Optimization

Jul 22, 2025
BL
Bingxin Liu
🏛️ Beijing Normal University | Phytium Technology Co., Ltd.

This work addresses key challenges in Profile-Guided Optimization (PGO): high sampling overhead, poor adaptability to dynamic inputs, and weak cross-architecture portability. We systematically survey and restructure the PGO technical landscape, proposing the first multi-dimensional classification framework for PGO—explicitly identifying three core research directions: low-overhead profiling, dynamic workload adaptation, and cross-architecture profile migration. Our approach unifies instrumentation- and sampling-based analysis, enabling compiler- and linker-time collaborative optimization in GCC and LLVM across heterogeneous targets including x86 and ARM. Empirical evaluation on standard benchmarks demonstrates an average performance improvement of 12.3%, while reducing profiling overhead to under 0.8%. These results significantly enhance the industrial deployability and generalization capability of PGO.

Addressing challenges like sampling overhead and cross-architecture portability.Enhancing performance through Profile Guided Optimization (PGO).Systematically categorizing PGO research by profiling methods and optimizations.

An Empirical Study: MEMS as a Static Performance Metric

May 12, 2025
LZ
Liwei Zhang
🏛️ Hangzhou Institute for Advanced Study | University of Chinese Academy of Sciences (UCAS) | Inria

Existing static performance analysis lacks lightweight, platform-agnostic metrics for compile-time performance estimation. Method: We propose memory access count (MEMS) as a static performance proxy, implemented via Clang AST rewriting and source-level automatic instrumentation to embed path-wise MEMS modeling and counting logic directly into the code. Contribution/Results: We systematically validate MEMS across ten classical algorithms. Experiments reveal— for the first time—that MEMS exhibits strong positive correlation with runtime performance *within* a given program across different execution paths, but weak correlation *across* distinct programs; this delineates MEMS’s practical applicability boundary. Crucially, MEMS incurs zero runtime overhead and requires no hardware-specific support, offering a portable, low-cost metric for static performance prediction. By bridging the gap between compile-time analysis and empirical performance behavior, MEMS significantly enhances the practicality of compiler optimizations and static performance modeling.

Assessing MEMS correlation with runtime performanceDeveloping a tool for automated MEMS countingEvaluating MEMS as a static performance metric

Latest Papers

What's happening recently
View more

Existing benchmarks inadequately capture the complexity of performance optimization in real-world codebases, often neglecting trade-offs between runtime and memory usage, measurement noise, and input variability. To address this gap, this work proposes SWE-Pro—the first repository-scale benchmark for performance optimization—constructed from expert-driven optimization cases across 102 open-source projects. SWE-Pro introduces multidimensional evaluation metrics, including parameterized testing, noise-aware measurement protocols, and Time-Weighted Memory Usage (TWMU), to holistically reflect practical engineering challenges. Experimental results demonstrate that current large language models exhibit limited effectiveness on this benchmark, achieving negligible runtime improvements and virtually no memory optimization, whereas expert solutions yield an average speedup of 15.5× and a 171.3× reduction in peak memory consumption, underscoring both the benchmark’s realism and its difficulty.

benchmarkexecution timeLarge Language Models

This work addresses the subtle microarchitectural performance inefficiencies often introduced by modern compiler optimizations, which can lead to significant yet overlooked performance losses. The authors propose a top-down differential analysis methodology that systematically identifies and categorizes the root causes of such optimization defects by integrating fine-grained microarchitectural performance counter sampling with cross-compiler (GCC/Clang) binary comparisons. Innovatively combining top-down microarchitectural analysis with differential testing, the approach further introduces a portable binary patching framework to precisely locate and rectify inefficient code segments. Empirical evaluation demonstrates that the method effectively uncovers substantial but commonly neglected performance discrepancies between GCC and Clang and successfully recovers performance through targeted binary patches.

binary performancecompiler optimizationdifferential analysis

This work addresses the challenge of memory inefficiencies—such as redundant allocations and suboptimal usage—in large-scale software systems, which often lead to significant resource waste and performance degradation. Existing optimization approaches lack end-to-end automation and struggle to scale to codebases exceeding hundreds of millions of lines. To overcome this, we propose MOA, a novel framework that integrates multi-agent large language models with performance profiling data. MOA employs three coordinated agents—Analyzer, Checker Generator, and Patcher—to automatically detect memory anti-patterns, synthesize static checkers, and generate state-machine-guided, semantics-preserving patches. Evaluated on OpenHarmony’s C/C++ codebase (>100 million lines), MOA identified 13 memory anti-patterns (9 previously unknown), pinpointed over 10,000 inefficiency instances, and produced 769 patches with a 92.5% expert acceptance rate, reducing heap memory usage by 42.2% and binary size by 10.6% on average.

codebase scalememory bloatmemory inefficiency

This work addresses the limited feature coverage in high-performance computing caused by hardware performance counters constrained by the number of simultaneously collectible metrics. To overcome this limitation without relying on hardware multiplexing, the authors propose a heuristic multi-run execution trace merging method that aligns and fuses counter data collected across multiple program executions. By analyzing MPI communication structures, timing patterns, and behavioral characteristics, the approach constructs a high-dimensional, unified synthetic trace that expands the effective feature space. This enriched representation enables the training of more comprehensive machine learning–based performance models. Experimental evaluation on the MareNostrum5 platform demonstrates that the merged counters retain high accuracy and significantly improve performance prediction for diverse kernel functions and real-world applications.

execution tracesfeature coveragehardware counters

Hot Scholars

LB

Luca Benini

ETH Zürich, Università di Bologna
Integrated CircuitsComputer ArchitectureEmbedded SystemsVLSI
HK

Hartmut Kaiser

Center of Computation and Technology at Louisiana State University
C++High Performance Parallel and Distributed ComputingRuntime SystemsCompiler Technologies
PD

Patrick Diehl

Los Alamos National Laboratory
Crack and fracture mechanicsPeridynamicsHPCHPX
MS

Marc Snir

University of Illinois at Urbana Chamapign
Parallel Computing
VJ

Vijay Janapa Reddi

Harvard University
Computer ArchitectureMachine Learning SystemsAutonomous Agents