performance optimization

Designs, implements, and evaluates measurement and profiling workflows and applies targeted changes to software, systems, or components to improve performance metrics such as latency, throughput, resource utilization, and scalability. Analyzes performance data to identify bottlenecks and applies algorithmic, architectural, code-level, or configuration-level optimizations (for example caching, concurrency, memory management, and runtime/compiler tuning).

performanceoptimization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.13
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$199K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Detecting Performance-Relevant Changes in Configurable Software Systems

Nov 21, 2025
SB
Sebastian Böhm
🏛️ Saarland University | Leipzig University

Detecting performance regressions in configurable software is costly, and configuration sampling often misses localized performance degradation. Method: This paper proposes ConfFLARE, a technique that combines data-flow dependency analysis with change-impact propagation tracking to identify code changes interacting—via data flow—with performance-sensitive code. It further integrates configuration-feature identification to automatically select the subset of performance-sensitive configurations most likely affected by each change. Contribution/Results: ConfFLARE eliminates the need for exhaustive configuration-based performance testing. In evaluations on synthetic and real-world systems, it reduces the number of required test configurations by 79% and 70%, respectively, while achieving near-complete coverage of performance regression cases. It precisely pinpoints relevant features and significantly improves both the efficiency and completeness of performance regression detection.

Detecting performance-relevant changes in configurable software systems efficientlyIdentifying performance regressions through data-flow interactions with critical codeReducing measurement costs by selecting relevant configurations for testing

gigiProfiler: Diagnosing Performance Issues by Uncovering Application Resource Bottlenecks

Jul 08, 2025
YH
Yigong Hu
🏛️ Boston University | University of Washington | University of California, Los Angeles

Modern software systems frequently exhibit application-level resource contention bottlenecks—such as blocking on custom application events—that evade detection by conventional performance profilers due to complex dependencies and bespoke resource management. To address this, we propose OmniResource Profiling, the first method to jointly leverage system-level metrics and application-level event-waiting relationships. It employs a lightweight LLM-assisted static analysis to automatically identify custom resources and cross-execution-trace runtime variable comparison for precise root-cause localization. Evaluated on 12 known performance issues across five real-world applications, OmniResource achieves 100% diagnostic accuracy and uncovers two previously undetected bottlenecks. Crucially, it requires no intrusive instrumentation, balancing high precision with practical deployability. This work delivers the first end-to-end solution for application-level resource contention analysis.

Diagnosing performance bottlenecks in complex applicationsIdentifying application-level resource contention issuesUncovering root causes of performance problems

Non-author engineers often struggle to associate performance bottlenecks with program semantics. Method: This paper proposes an interpretable optimization approach that jointly leverages runtime performance data and code semantics. It introduces CodeBERT—the first pre-trained code model—into performance profiling: fine-tuning it to generate fine-grained code summaries and aligning these with call-path-level performance profiles collected by Async Profiler for Java applications; hot paths and their semantic summaries are then co-visualized in a graphical interface. Contributions/Results: (1) We present the first semantic-augmented performance profiling framework built upon a pre-trained code model; (2) the approach significantly improves bottleneck interpretability and optimization guidance. Experiments across multiple Java benchmarks demonstrate that our system effectively reduces developers’ cognitive load, shortening average bottleneck localization time by 37.2%.

Enhancing profiler usability with deep learningInterpreting complex performance data from profilersLinking inefficiencies to program semantics automatically

When Should I Run My Application Benchmark?: Studying Cloud Performance Variability for the Case of Stream Processing Applications

Apr 16, 2025
SH
Soren Henning
🏛️ Dynatrace Research | Johannes Kepler University Linz

This study addresses the high variability and low reliability of performance benchmarking results for stream-processing applications in cloud environments. Over three months, we conducted a large-scale longitudinal empirical study across multiple geographic regions and heterogeneous hardware—including diverse CPU architectures. Leveraging Kubernetes-based automated deployment, high-frequency repeated benchmarking, and time-series statistical analysis, we systematically characterized end-to-end cloud performance variability for the first time at the application level. We discovered that variability exhibits statistically significant diurnal and weekly periodicity (amplitude ≤2.5%) and a coefficient of variation <3.7%—substantially lower than commonly assumed in industry. Moreover, infrastructure sharing incurs at most a 2.5-percentage-point loss in measurement precision. These findings demonstrate strong robustness across regions and CPU architectures, providing empirical evidence and methodological foundations for enhancing reproducibility and trustworthiness in cloud-native benchmarking.

Assess temporal effects on stream processing applicationsEvaluate benchmark result accuracy across cloud regionsQuantify cloud performance variability impact on benchmarks

Business process optimization remains challenging due to fragmented methodologies across process mining, predictive process monitoring, and process-aware recommendation—each operating in isolation without a unified theoretical foundation or integration framework. Method: This paper proposes a closed-loop optimization framework that systematically integrates Alpha algorithm/Inductive Miner for process discovery, LSTM/Transformer for runtime prediction, collaborative filtering/graph neural networks for action recommendation, and explainable AI (XAI) for interpretability—enabling automated bottleneck identification, anomaly forecasting, and prescriptive optimization from event logs. Contribution/Results: We establish the first unified conceptual boundary, evolutionary taxonomy, and synergy paradigm across the three domains; construct a comprehensive classification schema covering 120+ studies; clarify application scopes and standardized evaluation benchmarks; and deliver an industrially actionable methodology selection guide with validated deployment pathways.

Optimize business process performancePredict future process behaviorSupport data-driven decision-making

Latest Papers

What's happening recently
View more

Existing benchmarks inadequately capture the complexity of performance optimization in real-world codebases, often neglecting trade-offs between runtime and memory usage, measurement noise, and input variability. To address this gap, this work proposes SWE-Pro—the first repository-scale benchmark for performance optimization—constructed from expert-driven optimization cases across 102 open-source projects. SWE-Pro introduces multidimensional evaluation metrics, including parameterized testing, noise-aware measurement protocols, and Time-Weighted Memory Usage (TWMU), to holistically reflect practical engineering challenges. Experimental results demonstrate that current large language models exhibit limited effectiveness on this benchmark, achieving negligible runtime improvements and virtually no memory optimization, whereas expert solutions yield an average speedup of 15.5× and a 171.3× reduction in peak memory consumption, underscoring both the benchmark’s realism and its difficulty.

benchmarkexecution timeLarge Language Models

Existing evaluations of code-generating agents primarily emphasize functional correctness while overlooking their capacity for performance optimization in real-world scenarios. This work proposes the first end-to-end benchmark tailored to the full performance engineering lifecycle, requiring agents to achieve reproducible performance gains through profiling, diagnosis, code modification, and validation—all while preserving functional correctness. The framework introduces several innovations, including hidden correctness tests, verifiable speedup metrics, and trajectory auditing, integrated with program profiling, cross-layer bottleneck diagnosis, large model–agent collaboration, and an optimization-summary handoff strategy. Evaluation across seven long-horizon tasks and seven agent stacks reveals that optimization efficacy is highly workload-dependent, with no single dominant approach; relying solely on raw speedup ratios can lead to misleading conclusions, necessitating a holistic assessment that jointly considers correctness and reproducibility.

agentic tasksbenchmarkcoding agents

This work addresses the challenge of memory inefficiencies—such as redundant allocations and suboptimal usage—in large-scale software systems, which often lead to significant resource waste and performance degradation. Existing optimization approaches lack end-to-end automation and struggle to scale to codebases exceeding hundreds of millions of lines. To overcome this, we propose MOA, a novel framework that integrates multi-agent large language models with performance profiling data. MOA employs three coordinated agents—Analyzer, Checker Generator, and Patcher—to automatically detect memory anti-patterns, synthesize static checkers, and generate state-machine-guided, semantics-preserving patches. Evaluated on OpenHarmony’s C/C++ codebase (>100 million lines), MOA identified 13 memory anti-patterns (9 previously unknown), pinpointed over 10,000 inefficiency instances, and produced 769 patches with a 92.5% expert acceptance rate, reducing heap memory usage by 42.2% and binary size by 10.6% on average.

codebase scalememory bloatmemory inefficiency

Hot Scholars

ZY

Zhongliang Yang

Associate Professor, Beijing University of Posts and Telecommunications
AI SecurityFinTech
HQ

Huamin Qu

Chair Professor, Hong Kong University of Science and Technology
Data visualizationHuman-Computer InteractionExplainable AIE-Learning
CW

Chaozheng Wang

The Chinese University of Hong Kong
software engineeringartificial intelligence
BF

Bernd Finkbeiner

Professor of Computer Science, CISPA Helmholtz Center for Information Security
Reactive SystemsVerificationSynthesisTemporal Logic
BW

Benjamin Watson

Associate Professor of Computer Science, North Carolina State University
GamesComputer GraphicsHuman-Computer InterfacesVisualization