Score
Design, build, and maintain toolchains and build workflows that produce executable binaries for target architectures or platforms different from the build host, including configuring cross-compilers, linkers, assemblers, runtime libraries, sysroots, and platform-specific build scripts. Develop and validate cross-platform compilation processes by producing, testing, and debugging generated binaries on target hardware or emulators and integrating those toolchains into automated build environments.
ASIC development faces challenges in IP reuse and lacks integrated hardware-software co-verification and unified build infrastructure. Method: This paper introduces SoCMake—the first unified SoC build system supporting cross-compilation of Chisel/SystemRDL hardware descriptions with C/C++/assembly code. It integrates RTL generation, simulation, firmware compilation, and SoC configuration into a single workflow, overcoming the limited software compilation support of conventional hardware build tools. By deeply embedding SystemC, the RISC-V toolchain, and CMake’s extensibility framework, SoCMake enables automated, abstraction-level–aware co-building across hardware description → RTL → firmware. Contribution/Results: SoCMake has successfully accelerated iterative deployment of radiation-tolerant RISC-V SoCs in high-energy physics applications. After open-sourcing, it has become a de facto standard for generic SoC generation, reducing overall SoC development time by over 40% in empirical evaluations.
This work proposes PhantomRun, a novel framework that leverages large language models (LLMs) to automatically repair compilation failures in continuous integration (CI) pipelines for embedded open-source software—a domain often plagued by hardware dependencies, syntax errors, and build script issues that incur substantial debugging overhead. PhantomRun integrates build logs, source code, historical fixes, and error diagnostics to generate and validate repair patches. The framework incorporates an adapter layer to ensure compatibility with diverse CI platforms such as GitHub Actions and GitLab CI, as well as multiple build systems. Experimental evaluation on four widely used embedded software projects demonstrates that PhantomRun successfully resolves 45% of CI compilation failures, thereby establishing the effectiveness and practicality of LLMs in this challenging context.
This work addresses the frequent failures in continuous integration (CI) builds of embedded open-source software, which often stem from cross-compilation complexities, board-specific configurations, and toolchain constraints. These failures are compounded by heterogeneous, ephemeral build logs that are difficult to reuse. To tackle this challenge, the authors propose PhantomRun, a framework that enables standardized reproduction of historical failed builds through a build log abstraction layer, metadata standardization, containerized replay environments, and heterogeneous log parsing techniques. PhantomRun is the first system to support large-scale, controllable replay of failed embedded CI builds, offering a unified, machine-readable interface for build artifacts and metadata. Evaluated on 4,628 failed runs, PhantomRun successfully reconstructed 91.8% of the builds, with 98% preserving the original execution outcomes, demonstrating high reproducibility fidelity.
To address the trust gap between source code and binaries in untrusted build environments, this paper introduces *attestable builds*, a novel paradigm that ensures strong consistency from source snapshots to verifiable binaries—without modifying source code or build scripts. Leveraging trusted execution environments (TEEs) such as Intel SGX or AMD SEV alongside lightweight sandboxed containers, the approach integrates remote attestation protocols, formal modeling, and rigorous security verification to enable immediate, high-assurance source-to-binary mapping validation. Experimental evaluation demonstrates successful end-to-end builds of complex projects—including LLVM Clang—with zero source or script modifications. The system incurs only 42 seconds of startup latency and a 14% overhead in build time, while remaining resilient against powerful adversarial threats, including malicious builders and compromised infrastructure.
In software supply chain security, binaries built from identical source code on different platforms often exhibit bit-level discrepancies, rendering traditional byte-wise comparison ineffective for determining functional equivalence and detecting cross-build security risks. To address this, we propose a multi-level binary equivalence model and introduce the first clone-detection framework that jointly incorporates semantic- and behavioral-level equivalence reasoning—thereby overcoming the limitations of strict bitwise equality. Our approach integrates static analysis, bytecode parsing, and equivalence relation modeling to construct a verifiable, semi-synthetic benchmark and an automated equivalence decision system. Evaluated on 14,156 pairs of Java binaries, our method identifies bit-level differences in 26.49% of samples, yet accurately confirms their functional equivalence—demonstrating substantial improvements in trustworthiness and reliability of supply chain binary comparison.
This work addresses the challenge of detecting and repairing compilation errors caused by feature interactions in configurable C systems—errors that traditional compilers and existing variability-aware tools struggle to handle effectively. We present the first systematic exploration of leveraging foundation models for this task, proposing a variability-aware error detection and repair approach based on GPT-OSS-20B and Gemini 3 Pro. Our method is evaluated across synthetic systems, real-world GitHub commits, and mutation testing scenarios. Experimental results demonstrate that GPT-OSS-20B achieves 0.97 precision, 0.90 recall, and 0.94 accuracy on small-scale systems, successfully repairing over 70% of the errors. Notably, it also uncovers potential compilation defects in real Linux commits, offering a low-overhead, high-coverage alternative for variability-aware compilation.
This work addresses the unreliability of disassembly caused by the absence of compiler-intended semantic information in stripped binary executables. To overcome this limitation, the authors propose a novel lightweight metadata embedding mechanism that explicitly encodes critical semantics—such as code regions and memory boundaries—directly into the binary. This approach yields a decidable intermediate representation situated between raw binaries and source code. For the first time, it enables disassembly that is both decidable and recompilable, facilitating precise lifting to high-level intermediate representations. Experimental evaluation demonstrates that the embedded metadata incurs only 17% of the size overhead of DWARF debug information, introduces no runtime performance penalty, and successfully supports behavior-preserving binary lifting, instrumentation, and recompilation across a wide range of real-world C/C++ programs.
This work addresses the vulnerability of build system code to poisoning attacks, which pose a critical threat to software supply chain security. While existing tools primarily focus on application source code, they largely overlook the security of the build process itself. To bridge this gap, we propose a novel paradigm—“development-phase isolation”—that, for the first time, incorporates build scripts into the scope of security analysis. By leveraging information flow tracking and behavioral privilege modeling, our approach enables fine-grained monitoring of build-time code execution. We implement this methodology in a prototype tool, Foreman, which effectively detects anomalous and malicious behaviors within build scripts. In real-world evaluations, Foreman successfully identified the poisoned test files used in the recent XZ Utils supply chain attack, demonstrating both the efficacy and practicality of our approach.
Performance evaluation of WebAssembly lacks systematicity and transparency due to variations in runtimes, hardware, application domains, and benchmark diversity. To address this, this work proposes Wasure—the first modular and extensible benchmarking framework tailored for WebAssembly—that enables automated, cross-engine, cross-platform performance evaluation and dynamic program analysis. Using Wasure, we quantitatively demonstrate significant differences among widely used benchmark suites in terms of code coverage and execution behavior, thereby revealing the critical influence of benchmark diversity on evaluation outcomes. Our framework provides the community with a reproducible and transparent infrastructure for WebAssembly performance assessment.