Score
Designs, builds, and analyzes software code and its implementations with explicit focus on execution speed, memory and resource efficiency, and scalability. This includes selecting or redesigning algorithms and data structures, applying profiling-driven and low-level implementation changes, and using concurrency, caching, and compiler/runtime tuning to meet measurable performance targets.
Low-level systems (e.g., network stacks) resist efficient compile-time specialization for dynamic workloads and runtime environments due to high implementation complexity, difficulty in predicting optimal strategies, and continuously evolving conditions. This paper introduces Iridescent—the first online, automatic, measurement-driven JIT code specialization framework for low-level systems. Developers annotate only lightweight specialization points; the system autonomously explores and deploys optimal specialization strategies via JIT compilation and dynamic binary rewriting, guided by real-time performance feedback (e.g., latency, throughput), eliminating reliance on static compilation and manual modeling. Evaluated on real-world network stacks, Iridescent achieves significant performance improvements—up to 2.3× latency reduction and 1.8× throughput gain—while imposing minimal developer overhead. Crucially, it enables continuous optimization in response to dynamic workload shifts and heterogeneous hardware platforms, demonstrating robust adaptability across diverse deployment scenarios.
This work addresses key challenges in Profile-Guided Optimization (PGO): high sampling overhead, poor adaptability to dynamic inputs, and weak cross-architecture portability. We systematically survey and restructure the PGO technical landscape, proposing the first multi-dimensional classification framework for PGO—explicitly identifying three core research directions: low-overhead profiling, dynamic workload adaptation, and cross-architecture profile migration. Our approach unifies instrumentation- and sampling-based analysis, enabling compiler- and linker-time collaborative optimization in GCC and LLVM across heterogeneous targets including x86 and ARM. Empirical evaluation on standard benchmarks demonstrates an average performance improvement of 12.3%, while reducing profiling overhead to under 0.8%. These results significantly enhance the industrial deployability and generalization capability of PGO.
This work addresses the subtle microarchitectural performance inefficiencies often introduced by modern compiler optimizations, which can lead to significant yet overlooked performance losses. The authors propose a top-down differential analysis methodology that systematically identifies and categorizes the root causes of such optimization defects by integrating fine-grained microarchitectural performance counter sampling with cross-compiler (GCC/Clang) binary comparisons. Innovatively combining top-down microarchitectural analysis with differential testing, the approach further introduces a portable binary patching framework to precisely locate and rectify inefficient code segments. Empirical evaluation demonstrates that the method effectively uncovers substantial but commonly neglected performance discrepancies between GCC and Clang and successfully recovers performance through targeted binary patches.
Non-author engineers often struggle to associate performance bottlenecks with program semantics. Method: This paper proposes an interpretable optimization approach that jointly leverages runtime performance data and code semantics. It introduces CodeBERT—the first pre-trained code model—into performance profiling: fine-tuning it to generate fine-grained code summaries and aligning these with call-path-level performance profiles collected by Async Profiler for Java applications; hot paths and their semantic summaries are then co-visualized in a graphical interface. Contributions/Results: (1) We present the first semantic-augmented performance profiling framework built upon a pre-trained code model; (2) the approach significantly improves bottleneck interpretability and optimization guidance. Experiments across multiple Java benchmarks demonstrate that our system effectively reduces developers’ cognitive load, shortening average bottleneck localization time by 37.2%.
Existing benchmarks inadequately capture the complexity of performance optimization in real-world codebases, often neglecting trade-offs between runtime and memory usage, measurement noise, and input variability. To address this gap, this work proposes SWE-Pro—the first repository-scale benchmark for performance optimization—constructed from expert-driven optimization cases across 102 open-source projects. SWE-Pro introduces multidimensional evaluation metrics, including parameterized testing, noise-aware measurement protocols, and Time-Weighted Memory Usage (TWMU), to holistically reflect practical engineering challenges. Experimental results demonstrate that current large language models exhibit limited effectiveness on this benchmark, achieving negligible runtime improvements and virtually no memory optimization, whereas expert solutions yield an average speedup of 15.5× and a 171.3× reduction in peak memory consumption, underscoring both the benchmark’s realism and its difficulty.
This work addresses the limitations of existing code optimization techniques, which struggle to effectively handle dynamic languages and are often confined to single-level transformations, thereby failing to precisely identify performance bottlenecks. To overcome these challenges, we propose Optimo—a multi-level, pattern-aware code optimization framework powered by a Mixture-of-Prompts (MoP) architecture that, for the first time, integrates the mixture-of-experts paradigm into code optimization. Optimo employs differential profiling to pinpoint critical code structures and performs coordinated optimizations across four abstraction levels—from algorithms down to APIs. Experimental results demonstrate that Optimo achieves up to a 57.48% optimization success rate with a 3.97× speedup on human-written code, and a 42.42% success rate with a 13.51× speedup on LLM-generated code, significantly outperforming current baselines on the COFFE and Effibench benchmarks.
研究通过结合静态代码特征、动态执行轨迹和内核级资源数据,发现13种性能原型,并提出一个多信号回归检测框架,以提高软件性能分析和预测的准确性。
This work systematically evaluates the potential of large language models (LLMs) for automatic code optimization in high-performance computing (HPC), where traditional approaches often struggle to balance performance and correctness. The study introduces a novel methodology that leverages multi-level abstractions and goal-oriented prompting to guide LLMs in directly generating optimized C code. Evaluated on the PolyBench benchmark suite, this approach is compared against conventional auto-tuning frameworks that rely on schedule representations. Experimental results demonstrate that LLM-generated C code achieves superior performance and effectiveness, highlighting the critical influence of compiler optimization abstractions on LLM guidance. These findings establish a promising new direction toward verifiable, LLM-driven code optimization for HPC applications.
This work addresses the challenge of memory inefficiencies—such as redundant allocations and suboptimal usage—in large-scale software systems, which often lead to significant resource waste and performance degradation. Existing optimization approaches lack end-to-end automation and struggle to scale to codebases exceeding hundreds of millions of lines. To overcome this, we propose MOA, a novel framework that integrates multi-agent large language models with performance profiling data. MOA employs three coordinated agents—Analyzer, Checker Generator, and Patcher—to automatically detect memory anti-patterns, synthesize static checkers, and generate state-machine-guided, semantics-preserving patches. Evaluated on OpenHarmony’s C/C++ codebase (>100 million lines), MOA identified 13 memory anti-patterns (9 previously unknown), pinpointed over 10,000 inefficiency instances, and produced 769 patches with a 92.5% expert acceptance rate, reducing heap memory usage by 42.2% and binary size by 10.6% on average.
This study addresses the lack of systematic understanding regarding the energy impact of software refactoring under diverse workloads, particularly the absence of empirical analysis linking real-world refactoring practices to energy regressions. The authors construct a microbenchmark encompassing 68 refactoring types and a practical benchmark comprising 481 real refactoring commits from GitHub. Leveraging multi-workload scenarios and repeated paired energy measurements, they conduct the first large-scale investigation revealing that the energy effects of refactoring are highly workload-sensitive. Their findings demonstrate that refactoring type alone is insufficient to predict energy changes: 51.8% of refactorings in the microbenchmark and 7.5% in real projects induce statistically significant energy differences. Execution time explains energy variation only in controlled settings. Moreover, existing metrics and large language model–based approaches prove unreliable for detecting energy regressions.