Score
Designs and implements translators that lift compiled machine code (binaries) into a machine-independent intermediate representation (IR/p‑code), reconstructing control-flow and data-flow, normalizing calling conventions and idioms, and preserving program semantics across ISAs to enable further analysis or transformation.
This paper addresses engineering challenges in migrating large-scale cloud warehouse workloads from x86 to Arm instruction set architectures (ISAs), proposing a source-code recompilation–centric paradigm—distinct from binary translation. Drawing on nearly 40,000 real-world code commits at Google, we establish the first systematic task taxonomy for large-scale ISA migration. Our approach integrates static analysis, automated code refactoring, machine learning–assisted modifications, and CI pipeline monitoring to drive open-source ecosystem–based, full-stack software reconstruction. The methodology has been deployed internally at Google to automate x86-to-Arm migration across production systems, significantly improving efficiency while surfacing critical legacy bottlenecks. Key contributions include: (1) formalizing a recompilation-first framework for ISA migration; (2) introducing a principled, empirically grounded task classification system; and (3) empirically validating AI’s pivotal role in migration automation—providing an industry-reusable blueprint and opening new research directions in ISA migration for academia.
Binary program symbolic execution suffers from semantic distortion and implementation errors introduced during intermediate representation (IR) translation. Method: This paper proposes the first instruction-level symbolic execution framework directly grounded in formal ISA semantics (Rock/Sail), bypassing conventional IR abstractions by compiling machine-readable ISA specifications into SMT-solvable symbolic semantic models and integrating them into a binary analysis platform. Contributions/Results: (1) The first end-to-end automated pipeline from formal ISA semantics to symbolic execution; (2) Demonstrated scalability on RISC-V—modeling new instructions requires only a few hours; (3) Discovered five previously unknown ISA semantic implementation bugs in angr; (4) Achieved high-fidelity branch modeling and solving capability. The framework significantly improves the accuracy, trustworthiness, and development efficiency of binary symbolic execution.
This work investigates large language models’ (LLMs) capacity to comprehend compiler intermediate representations (IR), specifically evaluating their performance on four core tasks: control-flow graph (CFG) reconstruction, decompilation, code summarization, and execution reasoning. Through systematic multi-model benchmarking (GPT-4, LLaMA 3.1, Gemma 2, etc.), a curated structured IR dataset, task-specific prompt engineering, and fine-grained error attribution, the study provides the first empirical evidence of fundamental limitations in LLMs’ IR understanding—particularly in CFG reconstruction (accuracy <42%) and execution reasoning (error rate 68%). Methodologically, it introduces a dual-path enhancement paradigm: (1) IR-domain fine-tuning and (2) explicit control-flow modeling. Experimental results demonstrate that targeted fine-tuning improves task performance by up to 31.5%, establishing a foundational framework for advancing LLM-based IR analysis.
This work addresses the interoperability challenge between GCC and LLVM compiler intermediate representations (IRs), which stems from their semantic and structural differences. To bridge this gap, the authors propose IRIS-14B, the first large language model specifically designed for IR-to-IR translation. Built upon a 14-billion-parameter Transformer architecture, IRIS-14B leverages supervised fine-tuning to learn the mapping between GIMPLE and LLVM IR derived from the same C source code, enabling high-fidelity automatic translation. Experimental results demonstrate that IRIS-14B substantially outperforms existing open-source large models on real-world C programs and competitive programming tasks, achieving up to a 44-percentage-point improvement in accuracy. This study provides the first empirical validation of large language models as effective and feasible interoperability layers within neuro-symbolic hybrid compilation frameworks.
Traditional static binary translation struggles to reliably translate complete x86-64 programs to AArch64 in the absence of debugging information, source code, or assumptions about code layout. This work proposes a deterministic whole-program static translation approach that enumerates all possible instruction and data interpretations at the byte level, generating a complete set of candidate interpretations for each byte. By integrating these interpretations with code “tiles” automatically derived from high-level ISA semantics, the method synthesizes a fully functional target binary without relying on heuristic rules or runtime fallback mechanisms. For the first time, this enables static translation outputs that can be verified, tested, and signed ahead of execution. Experimental results demonstrate that the approach achieves performance on par with QEMU’s user-mode JIT on real-world benchmarks such as SPECint 2006, albeit at the cost of significantly larger output binary size.
This work addresses the challenge of reliable source-level binary patching in the absence of original source code and toolchains, where existing decompilers often produce outputs riddled with syntactic and semantic errors. To overcome this limitation, the authors propose a static patching framework that integrates decompilation with binary-aware recompilation. By leveraging information extracted directly from the original binary, the framework corrects semantic distortions in decompiled code and enables automated patch generation. The approach substantially improves recompilation correctness, fixing approximately 81% of erroneous functions produced by Hex-Rays, successfully patching 13 out of 14 real-world CVEs, and increasing user experiment success rates from 3.7% to 100%. Furthermore, it supports large-model-driven fully automated patching, demonstrating robust practical applicability.
The Instruction Set Architecture (ISA) is the contract between compilers and processors; proving this contract formally demands cross-level connection to existing mechanized compilers and hardware implementations. As an open, modular ISA gaining adoption across embedded, mobile, and cloud platforms, RISC-V makes a formally verified ISA specification particularly valuable. However, existing formal RISC-V specifications focus on hardware tooling rather than cross-level verification: they provide no machine-checked instruction-level properties and lack support for verifying this contract across levels. We address these limitations with a formal semantics of the RISC-V ISA in Rocq, built on Interaction Trees (ITrees). By leveraging ITree bisimulation and refinement, our semantics enables cross-level verification from compiler IR to hardware within a single framework. Our formalization covers a wide spectrum of RISC-V extensions. The correctness of individual instruction semantics is backed by machine-checked lemmas in Rocq. We further validate it by extracting an executable simulator that passes all standard RISC-V test suites. Three case studies demonstrate the effectiveness of our semantics for cross-level verification: first, we prove semantic equivalence via bisimulation between LLVM IR and RISC-V code on an array access pattern via Vellvm (LLVM ITree semantics); second, we apply translation validation to a specific instruction reordering for macro-operation fusion, distinguishing safe reorderings from those that break program-counter-relative addressing; third, we prove that a Kôika hardware ALU correctly implements all R-type integer operations (e.g., ADD, SUB, AND) against our ISA contract.
This work addresses the challenging problem of recovering original source code from stripped binary functions, a task where traditional decompilation typically yields only approximate pseudocode. The paper proposes a novel paradigm that replaces pseudocode generation with direct source code retrieval. By extracting anchors such as strings and constants from binaries, the method retrieves candidate functions from a source code corpus and constructs a multimodal representation incorporating assembly instructions, decompiled code, and metadata. A large language model (LLM) is then employed for semantic re-ranking of candidates. The approach integrates Ghidra-based static analysis with an inverted index system and introduces an iterative anchor refinement strategy. Evaluated on a high-quality tcpdump dataset, it achieves 95.2% instruction coverage, and attains 35.5% coverage on general-purpose GitHub repositories, demonstrating effectiveness in both ideal and noisy real-world scenarios.