LLMs as Idiomatic Decompilers: Recovering High-Level Code from x86-64 Assembly for Dart

📅 2026-04-02
📈 Citations: 0
Influential: 0
📄 PDF

career value

148K/year
🤖 AI Summary
This study addresses the open challenge in reverse engineering of decompiling x86-64 assembly into idiomatic modern high-level code, specifically Dart. The work proposes the first approach leveraging specialized small-scale large language models (4B/8B parameters), augmented with synthetically generated data and cross-lingual transfer from Swift to Dart, to recover high-quality Dart source code. Evaluated on a benchmark of 73 functions, the method achieves a CODEBLEU score of 71.3—approaching the performance of a 480B general-purpose model—and attains a compile@k5 rate of 79.4% on 34 real-world Dart functions, substantially outperforming baseline techniques. The results demonstrate that domain-specialized small models can produce semantically clear and idiomatic code, and reveal the existence of a model capacity threshold for effective cross-lingual transfer.

Technology Category

Application Category

📝 Abstract
Translating machine code into human-readable high-level languages is an open research problem in reverse engineering. Despite recent advancements in LLM-based decompilation to C, modern languages like Dart and Swift are unexplored. In this paper, we study the use of small specialized LLMs as an idiomatic decompiler for such languages. Additionally, we investigate the augmentation of training data using synthetic same-language examples, and compare it against adding human-written examples using related-language (Swift -> Dart). We apply CODEBLEU to evaluate the decompiled code readability and compile@k to measure the syntax correctness. Our experimental results show that on a 73-function Dart test dataset (representing diverse complexity levels), our 4B specialized model achieves 71.3 CODEBLEU (95% CI 65.5-77.1), approximately comparable to a ~480B code model (73.1; 67.4-78.8). On a subset of 34 natural Dart functions, it reaches compile@k5 = 79.4% (Wilson 95% CI 63.2-89.7), vs. 64.7% (47.9-78.5) for the base model; the difference is suggestive but not statistically significant at 0.05. Our results indicate that adding Swift training data helps at 8B but not at 4B, suggesting a capacity threshold for effective cross-lingual transfer. Our experimental results show that small specialized models can generate readable, idiomatic Dart with meaningful identifiers while using minimal compute.
Problem

Research questions and friction points this paper is trying to address.

decompilation
reverse engineering
Dart
LLM
machine code to high-level language
Innovation

Methods, ideas, or system contributions that make the work stand out.

idiomatic decompilation
specialized LLMs
cross-lingual transfer
CODEBLEU
compile@k