D-LiFT: Improving LLM-based Decompiler Backend via Code Quality-driven Fine-tuning

📅 2025-06-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Current LLM-based decompilers suffer from frequent syntactic/semantic errors, poor readability, error-prone quality improvements, and unreliable verification. To address these issues, this paper proposes an LLM-augmented system tailored for decompilation backends, adopting a precision-first reinforcement learning (RL) fine-tuning paradigm. We introduce D-SCORE—a novel multidimensional evaluation framework that integrates compilation success and symbolic execution–based verification into the reward loop, ensuring readability enhancements are applied exclusively to functionally correct code. Furthermore, we establish a precision–readability co-optimization principle. The system unifies Ghidra’s decompilation pipeline, LLM inference, RL optimization, compiler-based validation, and structured readability metrics. Evaluated on coreutils and util-linux benchmarks, our approach increases the number of functions achieving both high accuracy and high readability by 55.3% (measured by D-SCORE) over baseline LLM decompilers.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageComputer Vision: Large Vision Models

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
📝 Abstract
Decompilers, which reconstruct human-readable source code from binary executables, are vital to many security tasks. Yet, despite recent advances, their output often suffers from syntactic and semantic errors and remains difficult to read. Recently, with the advent of large language models (LLMs), researchers began to explore the potential of LLMs to refine decompiler output. Nevertheless, our study of these approaches reveals significant limitations, such as introducing new errors and relying on unreliable accuracy validation. In this paper, we present D-LiFT, an automated decompiler backend that harnesses and further trains LLMs to improve the quality of decompiled code via reinforcement learning (RL). Unlike prior work that overlooks preserving accuracy, D-LiFT adheres to a key principle for enhancing the quality of decompiled code: extit{preserving accuracy while improving readability}. Central to D-LiFT, we propose D-SCORE, an integrated quality assessment system to score the decompiled code from multiple aspects. In line with our principle, D-SCORE assigns low scores to any inaccurate output and only awards higher scores for readability to code that passes the accuracy check. Specifically, D-SCORE first verifies the syntactic and semantic correctness via the compiler and symbolic execution; only if a candidate is deemed accurate, it then evaluates readability using established metrics to compare the LLM output with the original decompiled code. The score will then be fed back to the LLM for fine-tuning. Our implementation, based on Ghidra and a range of LLMs, demonstrates significant improvements for the accurate decompiled code from the coreutils and util-linux projects. Compared to baseline LLMs without D-SCORE-driven fine-tuning, D-LiFT produces 55.3% more improved decompiled functions, as measured by D-SCORE.
Problem

Research questions and friction points this paper is trying to address.

Improving decompiled code quality via LLM fine-tuning
Ensuring accuracy while enhancing code readability
Validating decompilation correctness with integrated scoring
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses reinforcement learning for decompiler fine-tuning
Integrates D-SCORE for multi-aspect quality assessment
Ensures accuracy before improving code readability
🔎 Similar Papers
No similar papers found.