Score
Designs and builds reconstructed protocol specifications, decoders, and state-machine models that capture message formats, handshakes, and interaction semantics from observed traffic or opaque implementations. Analyzes binaries and firmware to recover binary protocol parsers, proprietary encodings or compression schemes, and to map parser reachability and state transitions for use in testing, emulation, or defensive tooling.
Existing static and dynamic analysis methods for inferring protocol state machines from implementations suffer from path explosion and insufficient code coverage, hindering accurate state-machine reconstruction. This paper introduces, for the first time, large language models (LLMs) to automate protocol state-machine inference. Our approach integrates retrieval-augmented generation (RAG), source-code semantic parsing, and structured prompt engineering to enable end-to-end identification, modeling, and formal representation of state machines directly from protocol source code. It overcomes classical analysis bottlenecks and supports precise characterization of cross-version state-machine differences. Evaluated on six mainstream protocol implementations, our method achieves over 90% state-machine recognition accuracy, improves fuzz-testing coverage by more than 20%, and successfully uncovers two zero-day vulnerabilities.
Traditional network protocol modeling via Markov chains heavily relies on manual expert knowledge, resulting in poor generalizability and scalability. This paper proposes the first end-to-end automated framework integrating large language models (LLMs) with state-aware fuzzing: an LLM automatically synthesizes executable state-transition models from protocol specifications and source code, then generates feedback-driven test sequences; program synthesis techniques further produce runnable fuzzers. The approach significantly lowers the barrier to protocol modeling and enhances cross-protocol transferability. Evaluated on three mainstream protocol implementations—TLS, DNS, and HTTP/2—the method discovered 12 previously unknown security vulnerabilities, all confirmed by developers. These findings validate the framework’s effectiveness and practical utility in real-world protocol security analysis.
Formal security analysis of closed-source encrypted applications (e.g., WhatsApp) remains intractable due to the absence of source code and cryptographic specifications. Method: We propose the first automated framework unifying functional correctness verification with microarchitectural side-channel resilience analysis. Leveraging Ghidra and an extended CryptoBAP, we perform binary-level reverse engineering to extract a formal model of WhatsApp’s encryption protocol—its first such model. We introduce hardware leakage contracts and integrate them with the DeepSec prover to jointly verify functional flaws and side-channel vulnerabilities. Contribution/Results: Our analysis uncovers previously unknown privacy violations invisible at the specification level—including contact leakage—and identifies a unlinkability attack against the BAC protocol. We formally verify forward secrecy, confirm susceptibility to cloning attacks, and expose deviations from the protocol specification. Crucially, we establish a reproducible, scalable, side-channel-aware formal analysis methodology for closed-source cryptographic software.
Addressing the “oracle absence” and “error attribution difficulty” challenges in network protocol parser verification, this paper proposes an LLM-driven framework for RFC semantic parsing and feedback-based oracle refinement. First, large language models automatically translate unstructured RFC text into formal message specifications. Second, an iterative, quasi-oracle is constructed to support specification-guided fuzz testing and cross-language (C/Python/Go) protocol implementation verification. Finally, vulnerabilities are precisely traced back to their originating RFC clauses. This work is the first to integrate LLM-based semantic understanding with dynamic oracle refinement. Evaluated on nine mainstream protocols, it discovers 69 vulnerabilities—36 of which have been confirmed—surpassing state-of-the-art approaches in both effectiveness and efficiency. It also demonstrates, for the first time, the feasibility of fully automated derivation of test oracles directly from natural-language protocol specifications.
Detecting and reproducing boundary-case vulnerabilities—especially those arising from state machine logic flaws in network protocols—remains challenging for conventional fuzzing due to inadequate coverage and poor reproducibility. Method: This paper proposes the first closed-loop approach integrating formal protocol specification inference, lightweight symbolic execution, and controllable vulnerability trace generation. It leverages SMT-driven state modeling, automatic synthesis of protocol interaction constraints, and automated proof-of-concept (PoC) generation to achieve end-to-end automation from vulnerability discovery to precise reproduction. Contribution/Results: Evaluated on 12 mainstream protocol stacks, the method discovers 17 previously unknown vulnerabilities—including 6 assigned CVEs—with an average reproduction time under 8 seconds and a false positive rate below 3%. It significantly improves accuracy, interpretability, and reproducibility in deep protocol vulnerability detection.
This work addresses the limitations of existing state-machine learning approaches, which rely on manually constructed input alphabets and often fail to cover anomalous or semantically invalid protocol messages. The paper presents the first fully automated framework for constructing protocol input alphabets: it leverages large language models to parse protocol message structures and applies structured mutation rules to generate input symbols encompassing both valid and invalid messages. To mitigate alphabet explosion, the approach incorporates a mini-batch incremental learning mechanism that reuses previously learned automata. Requiring no manual protocol knowledge, the method successfully reproduces known vulnerabilities and uncovers novel semantic flaws across multiple real-world protocol implementations, with several newly identified issues already acknowledged and patched by developers.
This study addresses the challenge of identifying undisclosed security fixes in commercial vehicle brake ECU firmware updates by proposing a methodology based on reverse engineering and differential binary analysis. The authors extract firmware images from the S12X architecture and, for the first time, apply QBinDiff to compare automotive ECU firmware patches, enabling precise localization of security-relevant code changes. Their analysis reveals that a recent safety recall was, in fact, a covert patch addressing a vulnerability in legacy protocol handling. The work successfully identifies critical security flaws and demonstrates that the update serves a dual purpose—providing both functional corrections and an unreported security patch—thereby establishing a novel paradigm for automotive firmware security auditing.
This work addresses the long-standing lack of systematic validation for processor specifications, which can lead to distorted program behavior and security vulnerabilities. It presents the first automated differential testing framework tailored for open-source SLEIGH specifications, automatically generating decodable instructions and initial execution states by parsing specification structures, and systematically validating them against multiple hardware reference implementations across architectures. Applied to x86-64 and AArch64, the approach uncovered 38,920 semantic discrepancies, identified 125 unique defects—many of which were subsequently fixed—and significantly improved specification fidelity. Furthermore, it exposed inconsistencies across vendor implementations and led to eight concrete recommendations, establishing a new paradigm for ensuring the reliability of instruction set architecture specifications.