protocol reverse engineering

Designs and builds reconstructed protocol specifications, decoders, and state-machine models that capture message formats, handshakes, and interaction semantics from observed traffic or opaque implementations. Analyzes binaries and firmware to recover binary protocol parsers, proprietary encodings or compression schemes, and to map parser reachability and state transitions for use in testing, emulation, or defensive tooling.

protocolreverseengineering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.51
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Inferring State Machine from the Protocol Implementation via Large Langeuage Model

May 01, 2024
HW
Haiyang Wei
🏛️ Nanjing University | QI-ANXIN Group | Southeast University

Existing static and dynamic analysis methods for inferring protocol state machines from implementations suffer from path explosion and insufficient code coverage, hindering accurate state-machine reconstruction. This paper introduces, for the first time, large language models (LLMs) to automate protocol state-machine inference. Our approach integrates retrieval-augmented generation (RAG), source-code semantic parsing, and structured prompt engineering to enable end-to-end identification, modeling, and formal representation of state machines directly from protocol source code. It overcomes classical analysis bottlenecks and supports precise characterization of cross-version state-machine differences. Evaluated on six mainstream protocol implementations, our method achieves over 90% state-machine recognition accuracy, improves fuzz-testing coverage by more than 20%, and successfully uncovers two zero-day vulnerabilities.

Complex Network CodeProtocol AnalysisState Machine Understanding

LLM-Assisted Model-Based Fuzzing of Protocol Implementations

Aug 03, 2025
CH
Changze Huang
🏛️ Key Lab of HCST | Peking University | NIO Inc.

Traditional network protocol modeling via Markov chains heavily relies on manual expert knowledge, resulting in poor generalizability and scalability. This paper proposes the first end-to-end automated framework integrating large language models (LLMs) with state-aware fuzzing: an LLM automatically synthesizes executable state-transition models from protocol specifications and source code, then generates feedback-driven test sequences; program synthesis techniques further produce runnable fuzzers. The approach significantly lowers the barrier to protocol modeling and enhances cross-protocol transferability. Evaluated on three mainstream protocol implementations—TLS, DNS, and HTTP/2—the method discovered 12 previously unknown security vulnerabilities, all confirmed by developers. These findings validate the framework’s effectiveness and practical utility in real-world protocol security analysis.

Automating network protocol testing using LLMsIdentifying vulnerabilities in protocol implementations efficientlyReducing manual effort in protocol model construction

Automated Side-Channel Analysis of Cryptographic Protocol Implementations

Nov 14, 2025
FN
Faezeh Nasrabadi
🏛️ CISPA Helmholtz Center for Information Security | Saarland University | KTH Royal Institute of Technology

Formal security analysis of closed-source encrypted applications (e.g., WhatsApp) remains intractable due to the absence of source code and cryptographic specifications. Method: We propose the first automated framework unifying functional correctness verification with microarchitectural side-channel resilience analysis. Leveraging Ghidra and an extended CryptoBAP, we perform binary-level reverse engineering to extract a formal model of WhatsApp’s encryption protocol—its first such model. We introduce hardware leakage contracts and integrate them with the DeepSec prover to jointly verify functional flaws and side-channel vulnerabilities. Contribution/Results: Our analysis uncovers previously unknown privacy violations invisible at the specification level—including contact leakage—and identifies a unlinkability attack against the BAC protocol. We formally verify forward secrecy, confirm susceptibility to cloning attacks, and expose deviations from the protocol specification. Crucially, we establish a reproducible, scalable, side-channel-aware formal analysis methodology for closed-source cryptographic software.

Analyzing protocol resilience against microarchitectural side-channel attacksExtracting formal models from binary implementations of cryptographic protocolsIdentifying security vulnerabilities invisible to specification-based analysis methods

Validating Network Protocol Parsers with Traceable RFC Document Interpretation

Apr 25, 2025
MZ
Mingwei Zheng
🏛️ Purdue University | Nanjing University

Addressing the “oracle absence” and “error attribution difficulty” challenges in network protocol parser verification, this paper proposes an LLM-driven framework for RFC semantic parsing and feedback-based oracle refinement. First, large language models automatically translate unstructured RFC text into formal message specifications. Second, an iterative, quasi-oracle is constructed to support specification-guided fuzz testing and cross-language (C/Python/Go) protocol implementation verification. Finally, vulnerabilities are precisely traced back to their originating RFC clauses. This work is the first to integrate LLM-based semantic understanding with dynamic oracle refinement. Evaluated on nine mainstream protocols, it discovers 69 vulnerabilities—36 of which have been confirmed—surpassing state-of-the-art approaches in both effectiveness and efficiency. It also demonstrates, for the first time, the feasibility of fully automated derivation of test oracles directly from natural-language protocol specifications.

Addressing oracle and traceability issues in protocol validationAutomating software validation via LLM-based specification translationValidating network protocol parsers using RFC documents

Formally Discovering and Reproducing Network Protocols Vulnerabilities

Mar 03, 2025
CC
Christophe Crochet
🏛️ Université catholique de Louvain

Detecting and reproducing boundary-case vulnerabilities—especially those arising from state machine logic flaws in network protocols—remains challenging for conventional fuzzing due to inadequate coverage and poor reproducibility. Method: This paper proposes the first closed-loop approach integrating formal protocol specification inference, lightweight symbolic execution, and controllable vulnerability trace generation. It leverages SMT-driven state modeling, automatic synthesis of protocol interaction constraints, and automated proof-of-concept (PoC) generation to achieve end-to-end automation from vulnerability discovery to precise reproduction. Contribution/Results: Evaluated on 12 mainstream protocol stacks, the method discovers 17 previously unknown vulnerabilities—including 6 assigned CVEs—with an average reproduction time under 8 seconds and a false positive rate below 3%. It significantly improves accuracy, interpretability, and reproducibility in deep protocol vulnerability detection.

Discover vulnerabilities in complex network protocolsReproduce vulnerabilities using attacker modelsTest single and multi-protocol interactions comprehensively

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing state-machine learning approaches, which rely on manually constructed input alphabets and often fail to cover anomalous or semantically invalid protocol messages. The paper presents the first fully automated framework for constructing protocol input alphabets: it leverages large language models to parse protocol message structures and applies structured mutation rules to generate input symbols encompassing both valid and invalid messages. To mitigate alphabet explosion, the approach incorporates a mini-batch incremental learning mechanism that reuses previously learned automata. Requiring no manual protocol knowledge, the method successfully reproduces known vulnerabilities and uncovers novel semantic flaws across multiple real-world protocol implementations, with several newly identified issues already acknowledged and patched by developers.

input alphabetprotocol implementationsemantic defects

This study addresses the challenge of identifying undisclosed security fixes in commercial vehicle brake ECU firmware updates by proposing a methodology based on reverse engineering and differential binary analysis. The authors extract firmware images from the S12X architecture and, for the first time, apply QBinDiff to compare automotive ECU firmware patches, enabling precise localization of security-relevant code changes. Their analysis reveals that a recent safety recall was, in fact, a covert patch addressing a vulnerability in legacy protocol handling. The work successfully identifies critical security flaws and demonstrates that the update serves a dual purpose—providing both functional corrections and an unreported security patch—thereby establishing a novel paradigm for automotive firmware security auditing.

Brake ECUfirmware updatelegacy protocol

This work addresses the long-standing lack of systematic validation for processor specifications, which can lead to distorted program behavior and security vulnerabilities. It presents the first automated differential testing framework tailored for open-source SLEIGH specifications, automatically generating decodable instructions and initial execution states by parsing specification structures, and systematically validating them against multiple hardware reference implementations across architectures. Applied to x86-64 and AArch64, the approach uncovered 38,920 semantic discrepancies, identified 125 unique defects—many of which were subsequently fixed—and significantly improved specification fidelity. Furthermore, it exposed inconsistencies across vendor implementations and led to eight concrete recommendations, establishing a new paradigm for ensuring the reliability of instruction set architecture specifications.

disassembleremulatorprocessor specification

Hot Scholars

XL

Xiapu Luo

The Hong Kong Polytechnic University
Mobile SecuritySmart ContractsNetwork SecurityBlockchain
MC

Mauro Conti

IEEE Fellow - Prof.@University of Padua - Wallenberg WASP Guest.Prof.@Örebro U.- Affiliate Prof.@UW
SecurityPrivacy
CP

Christof Paar

Max Planck Institute for Security and Privacy, Bochum
MH

Matthias Hollick

Professor of Computer Science, Technische Universität Darmstadt
Secure Mobile NetworkingNetwork SecurityMobile Networking
QW

Qin Wang

ETH Zurich
Domain AdaptationComputer Vision