Score
Designs and executes measurements, benchmarks, and analyses to evaluate and optimize the performance characteristics of programs written in Rust, including compile-time code generation effects (e.g., monomorphization), runtime optimizations (e.g., SIMD/vectorization), FFI boundary overheads, and performance tradeoffs of Rust’s safety model. Builds or configures profiling and benchmarking pipelines, proposes and validates code or compiler-level optimizations, and quantifies their impact on latency, throughput, and resource usage.
Despite Rust’s safety guarantees, rustc suffers from defects rooted in language-specific mechanisms—such as advanced traits, lifetime annotations, and unstable features—yet no systematic empirical study has characterized their root causes. Method: We conducted a large-scale manual root-cause analysis of 301 Rust-specific bugs reported between 2022 and 2024. We designed a Rust-semantics-aware defect taxonomy, mapped each bug to compiler phases (e.g., HIR, MIR), and evaluated the effectiveness of existing testing tools in detecting non-crashing semantic errors. Contribution/Results: Our study is the first empirical investigation revealing that over 70% of these defects originate in type checking and lifetime validation logic. We further find that current testing infrastructure exhibits significantly low detection rates for semantic errors—particularly those involving trait resolution and borrow-checker misbehavior. These findings provide reproducible, evidence-based guidance for improving rustc’s reliability and for designing targeted, semantics-aware test generation strategies.
This paper systematically investigates the core challenges in decompiling Rust binaries, identifying that Rust’s rich type system, aggressive compiler optimizations, and high-level abstractions—including generics, trait methods, and structured error handling—severely degrade decompilation fidelity. To address this, the authors introduce the first benchmark-driven, automated evaluation framework specifically designed for Rust decompilation, enabling quantitative assessment of control-flow reconstruction, variable naming, and type recovery across build configurations (e.g., debug vs. release). Experimental results reveal that monomorphization of generics and erasure of trait objects in release builds cause substantial loss of type information, while existing decompilers lack semantic awareness of Rust-specific constructs. The study provides an empirical foundation and concrete optimization directions for developing Rust-aware decompilation tools, highlighting critical gaps in current reverse-engineering infrastructure for memory-safe systems programming languages.
This work addresses key challenges in Profile-Guided Optimization (PGO): high sampling overhead, poor adaptability to dynamic inputs, and weak cross-architecture portability. We systematically survey and restructure the PGO technical landscape, proposing the first multi-dimensional classification framework for PGO—explicitly identifying three core research directions: low-overhead profiling, dynamic workload adaptation, and cross-architecture profile migration. Our approach unifies instrumentation- and sampling-based analysis, enabling compiler- and linker-time collaborative optimization in GCC and LLVM across heterogeneous targets including x86 and ARM. Empirical evaluation on standard benchmarks demonstrates an average performance improvement of 12.3%, while reducing profiling overhead to under 0.8%. These results significantly enhance the industrial deployability and generalization capability of PGO.
Rust lacks standardized benchmarks for scientific computing and high-performance computing (HPC), hindering its adoption and fair performance evaluation in these domains. Method: We present the first complete Rust port of the NAS Parallel Benchmarks (NPB), implementing all core kernels with full functionality. We systematically analyze the impact of Rust’s memory safety guarantees and ownership model on parallel performance, devise Rust-idiomatic parallelization strategies, and adopt Rayon as the parallel backend—comparing against optimized OpenMP-based Fortran and C++ implementations. Results: The sequential Rust variant achieves performance between Fortran and C++ (1.23% slower than Fortran, 5.59% faster than C++). While the Rayon-based parallel version does not yet surpass highly tuned OpenMP implementations, it demonstrates Rust’s feasibility for HPC workloads, engineering scalability, and potential as a principled, equitable platform for cross-language benchmarking—thereby filling a critical gap in standardized scientific computing evaluation for Rust.
Rust’s static memory safety guarantees can be violated during foreign function interface (FFI) interactions due to aliasing model incompatibilities—particularly with Tree Borrows—leading to undefined behavior (UB) that existing dynamic analysis tools like Miri cannot detect, creating a critical correctness gap in cross-language interoperability. Method: We conduct the first large-scale empirical study across 37 widely used Rust crates, combining Miri with the LLVM interpreter to enable cross-language cooperative analysis and systematically verify FFI call compliance under the Tree Borrows model. Contribution/Results: Our analysis uncovers 46 instances of UB or unexpected behavior—including in three high-download crates and one officially maintained Rust library—demonstrating that while Tree Borrows relaxes aliasing constraints, it exposes severe blind spots in current tooling for FFI contexts. This work establishes a novel methodology and an empirically grounded benchmark for verifying Rust’s cross-language memory safety.
This study addresses the challenge of verifying liveness properties in Rust asynchronous runtimes by proposing a lightweight, modular proof technique. Methodologically, it constructs a formal verification framework grounded in a model of the Rust language, integrating static analysis with a modular proof architecture. This approach pioneers a liveness verification paradigm for highly concurrent and heavily optimized libraries, overcoming traditional verification bottlenecks. Experimental results demonstrate that the proposed technique successfully verifies the eventual progress of multiple critical components, ensuring the reliable advancement of asynchronous tasks. Ultimately, this work provides a scalable pathway for formally guaranteeing low-level system infrastructure, significantly enhancing the reliability of the Rust asynchronous ecosystem.
Rust lacks a general-purpose dynamic analysis framework capable of supporting diverse runtime analyses. This work proposes DMIR, the first natively Rust-based, event-driven dynamic analysis infrastructure, which captures MIR-level semantics through compiler instrumentation and, for the first time, integrates high-level language features—such as ownership, types, and the memory model—into dynamic analysis. Runtime behaviors are exposed as structured event streams, enabling rich semantic introspection. Leveraging DMIR, we implement three classes of analysis tools: concolic execution, Rust-specific checkers, and control-flow tracing, demonstrating its expressiveness and practicality while maintaining acceptable runtime overhead.
This work addresses the subtle microarchitectural performance inefficiencies often introduced by modern compiler optimizations, which can lead to significant yet overlooked performance losses. The authors propose a top-down differential analysis methodology that systematically identifies and categorizes the root causes of such optimization defects by integrating fine-grained microarchitectural performance counter sampling with cross-compiler (GCC/Clang) binary comparisons. Innovatively combining top-down microarchitectural analysis with differential testing, the approach further introduces a portable binary patching framework to precisely locate and rectify inefficient code segments. Empirical evaluation demonstrates that the method effectively uncovers substantial but commonly neglected performance discrepancies between GCC and Clang and successfully recovers performance through targeted binary patches.
This work systematically evaluates the potential of large language models (LLMs) for automatic code optimization in high-performance computing (HPC), where traditional approaches often struggle to balance performance and correctness. The study introduces a novel methodology that leverages multi-level abstractions and goal-oriented prompting to guide LLMs in directly generating optimized C code. Evaluated on the PolyBench benchmark suite, this approach is compared against conventional auto-tuning frameworks that rely on schedule representations. Experimental results demonstrate that LLM-generated C code achieves superior performance and effectiveness, highlighting the critical influence of compiler optimization abstractions on LLM guidance. These findings establish a promising new direction toward verifiable, LLM-driven code optimization for HPC applications.
为解决C语言代码中的内存安全漏洞问题,提出TRACTOR基准测试,用于评估C到Rust的翻译工具的有效性和性能。