reverse engineer software

Analyze and reconstruct the internal structure, algorithms, control and data flow, and interfaces of software artifacts—including compiled binaries and obfuscated or minified JavaScript—by extracting code paths, data structures, file and protocol formats, and dependency relationships. Use those reconstructions to produce decompiled or deobfuscated representations, behavioral or protocol models, patches or compatibility shims, detection signatures/indicators, or to locate and characterize bugs and security issues.

reverseengineersoftware

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.37
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$215K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

ReF Decompile: Relabeling and Function Call Enhanced Decompile

Feb 17, 2025
YF
Yunlong Feng
🏛️ Harbin Institute of Technology | East China Normal University

Existing end-to-end decompilation methods struggle to accurately recover control-flow structures and variable semantics, limiting logical reconstruction fidelity. This paper proposes an end-to-end large language model framework tailored for binary decompilation. Its core contributions are: (1) a *Relabelling* strategy that replaces jump addresses with semantic labels to explicitly model control-flow graphs; and (2) a *Function Call* strategy that infers variable types and restores missing symbolic information via call-site context. The framework tightly integrates instruction-level relabelling, function-call modeling, and binary symbol extraction. Evaluated on the Humaneval-Decompile benchmark, it achieves 61.43% functional correctness—significantly surpassing prior state-of-the-art methods. The approach robustly supports downstream security tasks, including vulnerability discovery, malware analysis, and legacy system migration.

Decompilation of low-level codeInferring variable types accuratelyPreserving control flow clarity

This work addresses the challenge of maintaining up-to-date architectural documentation in microservice systems, which is exacerbated by polyglot implementations, multiple repositories, and rapid independent evolution. Existing static refactoring approaches are often limited to single-repository settings or homogeneous technology stacks. To overcome these limitations, we propose a distributed static architecture reconstruction framework that supports multi-language and multi-repository environments. The framework employs pluggable extractor modules for language-specific analysis and introduces mechanisms for cross-repository data propagation and fusion, enabling seamless interoperability with existing static analysis tools. To the best of our knowledge, this is the first framework to enable distributed, collaborative architecture reconstruction, significantly enhancing the scalability and usability of automated documentation generation and maintenance in complex microservice ecosystems.

architecture reconstructionmicroservicemulti-repository

Reimagining Disassembly Interfaces with Visualization: Combining Instruction Tracing and Control Flow with DisViz

Oct 21, 2025
SH
Shadmaan Hye
🏛️ SCI Institute | Lawrence Livermore National Laboratory

Binary disassembly analysis suffers from ambiguous source-to-instruction mapping and difficulty in jointly preserving execution order and control flow. To address this, we propose DisViz—a performance-analysis-oriented, interactive disassembly visualization tool. Its core contributions are threefold: (1) a basic-block–based instruction layout that explicitly preserves execution order while intuitively revealing control structures (e.g., loops); (2) block-level minimaps to enhance contextual awareness and navigation in large-scale disassembly; and (3) integrated instruction tracing, control-flow graph visualization, and dynamic source-code correlation, enabling bidirectional, web-based navigation between source and disassembly. An empirical evaluation with ten domain experts from diverse institutions demonstrates that DisViz significantly improves both accuracy in identifying compiler optimization behaviors and overall analysis efficiency—validating its effectiveness for understanding compilation transformations and their performance implications.

Addressing challenges in mapping binary instructions to source codeImproving developer comprehension of compiler optimizations in binariesVisualizing disassembly code with execution order and control flow

Latest Papers

What's happening recently
View more

This work addresses the challenge of reliable source-level binary patching in the absence of original source code and toolchains, where existing decompilers often produce outputs riddled with syntactic and semantic errors. To overcome this limitation, the authors propose a static patching framework that integrates decompilation with binary-aware recompilation. By leveraging information extracted directly from the original binary, the framework corrects semantic distortions in decompiled code and enables automated patch generation. The approach substantially improves recompilation correctness, fixing approximately 81% of erroneous functions produced by Hex-Rays, successfully patching 13 out of 14 real-world CVEs, and increasing user experiment success rates from 3.7% to 100%. Furthermore, it supports large-model-driven fully automated patching, demonstrating robust practical applicability.

binary patchingdecompilationrecompilation

This work addresses the challenging problem of recovering original source code from stripped binary functions, a task where traditional decompilation typically yields only approximate pseudocode. The paper proposes a novel paradigm that replaces pseudocode generation with direct source code retrieval. By extracting anchors such as strings and constants from binaries, the method retrieves candidate functions from a source code corpus and constructs a multimodal representation incorporating assembly instructions, decompiled code, and metadata. A large language model (LLM) is then employed for semantic re-ranking of candidates. The approach integrates Ghidra-based static analysis with an inverted index system and introduces an iterative anchor refinement strategy. Evaluated on a high-quality tcpdump dataset, it achieves 95.2% instruction coverage, and attains 35.5% coverage on general-purpose GitHub repositories, demonstrating effectiveness in both ideal and noisy real-world scenarios.

binary functionsbinary-to-source matchingreverse engineering

Virtualization-obfuscated binary code is typically large and structurally complex, exceeding the input length limits of large language models (LLMs) and lacking annotated data, which hinders direct application in code analysis. To address this challenge, this work proposes a structure-role-oriented decomposition and automatic labeling paradigm: it employs static analysis to partition obfuscated code into maximally sized, semantically coherent units that conform to LLM input constraints, and automatically annotates each unit based on its structural role within the control flow graph. This approach enables the construction of a scalable dataset for both training and inference. Evaluation on real-world virtualization-obfuscated binaries demonstrates that the prototype system achieves efficient and accurate analysis, effectively overcoming the dual bottlenecks of input length limitations and data scarcity that currently impede LLM-based approaches in this domain.

input size limitslarge-scale labeled dataLLM-based analysis

Existing code debloating research relies on proxy metrics such as test coverage or code size, lacking evaluation grounded in actual program behavior. This work presents the first unified evaluation framework benchmarked against real-world execution semantics, integrating dynamic and static analysis to systematically reassess eight state-of-the-art debloating tools across source code, intermediate representations, and binary levels. The study reveals that dynamic approaches can erroneously remove up to 94% of code that should be retained, while static methods suffer from high false retention rates due to over-approximation—sometimes even introducing new, unintended function variants—that critically compromise program correctness and security. These findings expose systematic biases in current debloating techniques and establish an empirical foundation for future tool development.

application-levelcode debloatingevaluation methodology

Hot Scholars

YZ

Yanjie Zhao

Huazhong University of Science and Technology
Software EngineeringSoftware Security
GB

Guangdong Bai

Associate Professor of The University of Queensland
System SecuritySoftware SecurityTrustworthy AIPrivacy Compliance
DL

David Lo

Professor of Computer Science, Singapore Management University
AI4SESoftware AnalyticsSE4AISoftware Maintenance
LW

Liuhuo Wan

PhD, University of Queensland
cyber securitysoftware engineering
JL

Junchao Li

Ph.D., Mechanical Engineering, University of Iowa
Machine LearningReinforcement learningFormal methodsPath Planning