An Automated Framework for Input Alphabet Construction in Stateful Protocol Implementation Learning

📅 2026-06-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing state-machine learning approaches, which rely on manually constructed input alphabets and often fail to cover anomalous or semantically invalid protocol messages. The paper presents the first fully automated framework for constructing protocol input alphabets: it leverages large language models to parse protocol message structures and applies structured mutation rules to generate input symbols encompassing both valid and invalid messages. To mitigate alphabet explosion, the approach incorporates a mini-batch incremental learning mechanism that reuses previously learned automata. Requiring no manual protocol knowledge, the method successfully reproduces known vulnerabilities and uncovers novel semantic flaws across multiple real-world protocol implementations, with several newly identified issues already acknowledged and patched by developers.
📝 Abstract
As a prevalent analytical technique for stateful protocol implementations, state machine learning suffers from a core bottleneck stemming from handcrafted input alphabets. Manual alphabet definition inherently limits the completeness of input exploration, making it difficult to capture anomalous non-conformant messages and consequently missing latent semantic defects. In this paper, we target automatic input alphabet generation to break the above limitation for state machine learning. We adopt large language models to parse protocol message layouts and produce candidate input symbols following structured mutation rules, which automatically covers valid and invalid message spaces and eliminates reliance on manual protocol expertise. Considering the rising overhead brought by continuously growing alphabets, we introduce a mini-batch incremental learning strategy to reuse existing learned automata when incorporating new alphabet entries. Comprehensive experiments on practical protocol stacks indicate our approach can reproduce existing security vulnerabilities and identify novel semantic bugs. A subset of these newly discovered issues has been confirmed and patched by developers, proving the practicability and effectiveness of our proposed method.
Problem

Research questions and friction points this paper is trying to address.

stateful protocol
input alphabet
state machine learning
semantic defects
protocol implementation
Innovation

Methods, ideas, or system contributions that make the work stand out.

automated alphabet construction
stateful protocol learning
large language models
structured mutation
incremental automata learning