c++ performance optimization

Designs, implements, and analyzes performance improvements across software layers and runtimes, covering low‑level C/C++ and C/C++ toolchains, Python performance, runtime and systems tuning, and front‑end/UI and web/API latency and throughput. Work includes profiling and benchmarking to identify CPU, memory, I/O, and concurrency bottlenecks and applying algorithmic changes, caching, parallelism, compiler/build and low‑level optimizations, and measures for performance portability and accelerator (e.g., TPU) tuning.

c++performanceoptimization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.5
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$213K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

From Profiling to Optimization: Unveiling the Profile Guided Optimization

Jul 22, 2025
BL
Bingxin Liu
🏛️ Beijing Normal University | Phytium Technology Co., Ltd.

This work addresses key challenges in Profile-Guided Optimization (PGO): high sampling overhead, poor adaptability to dynamic inputs, and weak cross-architecture portability. We systematically survey and restructure the PGO technical landscape, proposing the first multi-dimensional classification framework for PGO—explicitly identifying three core research directions: low-overhead profiling, dynamic workload adaptation, and cross-architecture profile migration. Our approach unifies instrumentation- and sampling-based analysis, enabling compiler- and linker-time collaborative optimization in GCC and LLVM across heterogeneous targets including x86 and ARM. Empirical evaluation on standard benchmarks demonstrates an average performance improvement of 12.3%, while reducing profiling overhead to under 0.8%. These results significantly enhance the industrial deployability and generalization capability of PGO.

Addressing challenges like sampling overhead and cross-architecture portability.Enhancing performance through Profile Guided Optimization (PGO).Systematically categorizing PGO research by profiling methods and optimizations.

This work addresses the subtle microarchitectural performance inefficiencies often introduced by modern compiler optimizations, which can lead to significant yet overlooked performance losses. The authors propose a top-down differential analysis methodology that systematically identifies and categorizes the root causes of such optimization defects by integrating fine-grained microarchitectural performance counter sampling with cross-compiler (GCC/Clang) binary comparisons. Innovatively combining top-down microarchitectural analysis with differential testing, the approach further introduces a portable binary patching framework to precisely locate and rectify inefficient code segments. Empirical evaluation demonstrates that the method effectively uncovers substantial but commonly neglected performance discrepancies between GCC and Clang and successfully recovers performance through targeted binary patches.

binary performancecompiler optimizationdifferential analysis

Towards Online Code Specialization of Systems

Jan 20, 2025
VA
Vaastav Anand
🏛️ Max Planck Institute for Software Systems

Low-level systems (e.g., network stacks) resist efficient compile-time specialization for dynamic workloads and runtime environments due to high implementation complexity, difficulty in predicting optimal strategies, and continuously evolving conditions. This paper introduces Iridescent—the first online, automatic, measurement-driven JIT code specialization framework for low-level systems. Developers annotate only lightweight specialization points; the system autonomously explores and deploys optimal specialization strategies via JIT compilation and dynamic binary rewriting, guided by real-time performance feedback (e.g., latency, throughput), eliminating reliance on static compilation and manual modeling. Evaluated on real-world network stacks, Iridescent achieves significant performance improvements—up to 2.3× latency reduction and 1.8× throughput gain—while imposing minimal developer overhead. Crucially, it enables continuous optimization in response to dynamic workload shifts and heterogeneous hardware platforms, demonstrating robust adaptability across diverse deployment scenarios.

Low-level system optimizationPerformance enhancementRuntime adaptation

Existing benchmarks inadequately capture the complexity of performance optimization in real-world codebases, often neglecting trade-offs between runtime and memory usage, measurement noise, and input variability. To address this gap, this work proposes SWE-Pro—the first repository-scale benchmark for performance optimization—constructed from expert-driven optimization cases across 102 open-source projects. SWE-Pro introduces multidimensional evaluation metrics, including parameterized testing, noise-aware measurement protocols, and Time-Weighted Memory Usage (TWMU), to holistically reflect practical engineering challenges. Experimental results demonstrate that current large language models exhibit limited effectiveness on this benchmark, achieving negligible runtime improvements and virtually no memory optimization, whereas expert solutions yield an average speedup of 15.5× and a 171.3× reduction in peak memory consumption, underscoring both the benchmark’s realism and its difficulty.

benchmarkexecution timeLarge Language Models

THAPI: Tracing Heterogeneous APIs

Mar 22, 2025
SB
Solomon Bekele
🏛️ Argonne National Laboratory | University De Los Andes

Programming heterogeneous exascale HPC systems is hindered by the proliferation of complex, incompatible programming models (e.g., CUDA, SYCL, OpenMP) and the lack of traceability across CPU/GPU execution contexts. Method: We propose the first semantic-aware, full-stack API tracing framework built upon LTTng kernel tracing, integrating user-space dynamic instrumentation with multi-model API signature parsing to capture fine-grained, low-overhead, configurable API call chains across hardware and programming abstractions. Contribution/Results: Unlike conventional tracers logging only function names and timestamps, our framework enables cross-vendor, cross-abstraction behavioral correlation and end-to-end call-chain reconstruction in real HPC applications. It accurately identifies cross-model performance bottlenecks and implementation flaws, improving debugging efficiency by over 3× and significantly enhancing portability and debuggability of heterogeneous programming models.

Capturing detailed API calls across HPC software stack layersDebugging and optimizing performance in heterogeneous computing environmentsUnderstanding interactions between multiple programming models in HPC systems

Latest Papers

What's happening recently
View more

Existing benchmarks for C++ performance repair are often based on competitive programming code or focus on other languages, lacking executable, real-world evaluation data. This work proposes CppPerf-Mine, an automated pipeline that mines genuine performance-improving commits from GitHub repositories and constructs CppPerf-DB—the first large-scale, multi-file, reproducible dataset of C++ performance enhancements. The pipeline integrates structured filtering, large language model–based classification, and Dockerized build-and-test validation. CppPerf-DB comprises 347 human-verified patches spanning 42 mature projects, with 39% involving modifications across multiple files. Preliminary evaluation shows that OpenHands, a state-of-the-art tool, successfully repairs only 13.5% of the cases, underscoring the dataset’s critical value for advancing repository-level performance repair research.

automated program repairbenchmark datasetC++ performance bugs

This study addresses the absence of open standards for CPU pipeline visualization tools and the difficulty in localizing performance bottlenecks. To this end, it proposes an open-source event stream format alongside Catscan, an interactive viewer. Methodologically, this work introduces a structured event stream based on transactional relationships, integrating typed event modeling, persistent highlighting techniques, and domain-specific search algorithms to enable microarchitectural trace analysis from symptoms down to individual instructions. Furthermore, it supports resource-oriented views synchronized with comparative trace alignment. By successfully reproducing industry-grade debugging workflows, this project provides the community with production-validated microarchitectural visualization infrastructure.

CPU performance simulationmicroarchitecture debuggingopen-source tooling

Existing benchmarks for execution-time optimization patches primarily target Python, C++, or .NET, lacking configurable, reproducible solutions tailored to Java. This work proposes JETO-Mine, the first framework for automatically mining and validating Java performance patches with customizable filtering and statistically rigorous validation. JETO-Mine employs a three-stage pipeline integrating static analysis, LLM-driven issue categorization, Docker-based dynamic testing, and significance testing to construct JETO-Bench—a benchmark comprising 660 candidate patches and 91 manually verified effective ones. Experimental evaluation demonstrates that JETO-Bench effectively assesses patch generation tools (e.g., OpenHands achieves a 14.3% repair success rate) and reveals a widespread absence of performance validation tests in Java projects.

benchmarkexecution time improvementJava

Performance bugs in just-in-time (JIT) compilers can severely degrade program efficiency, yet systematic studies and automated detection approaches have been lacking. This work presents the first empirical analysis of 191 real-world JIT performance bugs, uncovering their triggering mechanisms, manifestation patterns, and root causes. To address this challenge, the authors propose a lightweight, hierarchical differential performance testing method that integrates test case prioritization with an automated deduplication and filtering mechanism, substantially improving both detection efficiency and precision. The implemented tool, Jittery, identified 12 previously unknown performance bugs in Oracle HotSpot and GraalVM—11 confirmed and 6 already fixed—while reducing testing time by 92.40%.

bug detectionempirical studyJIT compiler

Hot Scholars

LB

Luca Benini

ETH Zürich, Università di Bologna
Integrated CircuitsComputer ArchitectureEmbedded SystemsVLSI
WX

Wayne Xin Zhao

Professor, Renmin University of China
Recommender SystemNatural Language ProcessingLarge Language Model
WL

Weiwen Liu

Associate Professor, Shanghai Jiao Tong University
large language modelsAI agentsrecommender systems
JR

Ji-Rong Wen

Gaoling School of Artificial Intelligence, Renmin University of China
Large Language ModelWeb SearchInformation RetrievalMachine Learning
TH

Torsten Hoefler

Professor of Computer Science at ETH Zurich
High Performance ComputingDeep LearningNetworkingMessage Passing Interface