Score
Design and implement systems that ingest disassembled binary code and parse it into structured, machine-readable representations: detect binary/container formats and ISAs, extract instruction streams, metadata, and relocation information, and normalize instructions into a common intermediate representation. Build analyses and tooling that resolve calling conventions and reconstruct control-flow and data-flow so the parsed output can be analyzed, transformed, or recompiled.
Binary disassembly analysis suffers from ambiguous source-to-instruction mapping and difficulty in jointly preserving execution order and control flow. To address this, we propose DisViz—a performance-analysis-oriented, interactive disassembly visualization tool. Its core contributions are threefold: (1) a basic-block–based instruction layout that explicitly preserves execution order while intuitively revealing control structures (e.g., loops); (2) block-level minimaps to enhance contextual awareness and navigation in large-scale disassembly; and (3) integrated instruction tracing, control-flow graph visualization, and dynamic source-code correlation, enabling bidirectional, web-based navigation between source and disassembly. An empirical evaluation with ten domain experts from diverse institutions demonstrates that DisViz significantly improves both accuracy in identifying compiler optimization behaviors and overall analysis efficiency—validating its effectiveness for understanding compilation transformations and their performance implications.
This work addresses the unreliability of disassembly caused by the absence of compiler-intended semantic information in stripped binary executables. To overcome this limitation, the authors propose a novel lightweight metadata embedding mechanism that explicitly encodes critical semantics—such as code regions and memory boundaries—directly into the binary. This approach yields a decidable intermediate representation situated between raw binaries and source code. For the first time, it enables disassembly that is both decidable and recompilable, facilitating precise lifting to high-level intermediate representations. Experimental evaluation demonstrates that the embedded metadata incurs only 17% of the size overhead of DWARF debug information, introduces no runtime performance penalty, and successfully supports behavior-preserving binary lifting, instrumentation, and recompilation across a wide range of real-world C/C++ programs.
Existing binary disassemblers are commonly evaluated using source-code–based validation, an assumption that fails in real-world scenarios lacking source code—such as security analysis of closed-source software. Method: This paper introduces TraceBin, the first systematic, source-free methodology for evaluating disassembly errors solely from binary executables. TraceBin combines dynamic execution tracing with fine-grained binary analysis to precisely identify control-flow–related errors that directly impact security-critical tasks—including static instrumentation, hardening, and automated patching. Contribution/Results: TraceBin discovers novel disassembly flaws in non-C/C++ and closed-source binaries—previously unreported. Experimental evaluation across mainstream disassemblers reveals numerous known and previously unknown errors, significantly improving the reliability and practicality of security analyses on closed-source binaries.
This work addresses the challenge of reliable source-level binary patching in the absence of original source code and toolchains, where existing decompilers often produce outputs riddled with syntactic and semantic errors. To overcome this limitation, the authors propose a static patching framework that integrates decompilation with binary-aware recompilation. By leveraging information extracted directly from the original binary, the framework corrects semantic distortions in decompiled code and enables automated patch generation. The approach substantially improves recompilation correctness, fixing approximately 81% of erroneous functions produced by Hex-Rays, successfully patching 13 out of 14 real-world CVEs, and increasing user experiment success rates from 3.7% to 100%. Furthermore, it supports large-model-driven fully automated patching, demonstrating robust practical applicability.
Existing decompilation evaluations predominantly rely on syntactic similarity or isolated readability metrics, which inadequately capture the practical reusability of recovered code. To address this limitation, this work proposes a three-dimensional evaluation paradigm centered on reusability—encompassing readability, recompilability, and functionality—and introduces DEBENCH, the first automated multidimensional benchmark comprising 240 atomic functions and 640 binary samples. Leveraging LLM-as-judge for readability scoring, URAF fine-grained metrics, 50-round iterative compilation repair, and Frida-driven multilevel dynamic differential tracing, the study systematically uncovers significant discrepancies across evaluation dimensions: only 1.2% of outputs from the best decompiler–LLM combination achieve full functional equivalence; Clang-generated code exhibits 2.6× higher functionality than GCC’s; and functional recovery capability varies by up to 20× across decompilers. The analysis further identifies three dominant failure modes, including type system collapse.
This work addresses a critical limitation in existing binary code representation learning methods, which typically overlook instruction-level alignment information and thus fail to effectively leverage fine-grained supervisory signals from compiler debug information. To overcome this, the paper introduces the first approach that explicitly models instruction alignment as an auxiliary training objective. By employing multi-task learning, the method jointly optimizes function-level embeddings and instruction alignment, using debug information to construct precise alignment supervision signals. Experimental results demonstrate that this approach significantly improves accuracy in binary code similarity retrieval, enhances the model’s discriminative power and semantic understanding, and reveals a strong correlation between instruction alignment and the quality of function representations. Consequently, it establishes a more interpretable and precise framework for binary code representation learning.
This work addresses the challenging problem of recovering original source code from stripped binary functions, a task where traditional decompilation typically yields only approximate pseudocode. The paper proposes a novel paradigm that replaces pseudocode generation with direct source code retrieval. By extracting anchors such as strings and constants from binaries, the method retrieves candidate functions from a source code corpus and constructs a multimodal representation incorporating assembly instructions, decompiled code, and metadata. A large language model (LLM) is then employed for semantic re-ranking of candidates. The approach integrates Ghidra-based static analysis with an inverted index system and introduces an iterative anchor refinement strategy. Evaluated on a high-quality tcpdump dataset, it achieves 95.2% instruction coverage, and attains 35.5% coverage on general-purpose GitHub repositories, demonstrating effectiveness in both ideal and noisy real-world scenarios.
This study addresses the open challenge in reverse engineering of decompiling x86-64 assembly into idiomatic modern high-level code, specifically Dart. The work proposes the first approach leveraging specialized small-scale large language models (4B/8B parameters), augmented with synthetically generated data and cross-lingual transfer from Swift to Dart, to recover high-quality Dart source code. Evaluated on a benchmark of 73 functions, the method achieves a CODEBLEU score of 71.3—approaching the performance of a 480B general-purpose model—and attains a compile@k5 rate of 79.4% on 34 real-world Dart functions, substantially outperforming baseline techniques. The results demonstrate that domain-specialized small models can produce semantically clear and idiomatic code, and reveal the existence of a model capacity threshold for effective cross-lingual transfer.