Score
Designs and builds finite-state machine (FSM) models that represent interaction protocols or behaviors by segmenting recorded interaction traces into states, assigning state labels, and inferring transitions and their triggering conditions. Produces interpretable interaction graphs or state diagrams and validates those FSMs against observed traces to ensure they explain or predict interaction sequences.
This work addresses the challenge of generating reliable system models directly from natural language requirements by proposing a novel framework that leverages large language models (LLMs) for the automatic generation and repair of finite state machines (FSMs). It introduces, for the first time, the use of LLMs in FSM construction and innovatively integrates FSM mutation analysis with automated test generation to establish an expert feedback–driven optimization mechanism. Experimental validation on synthetic datasets using GPT-4 demonstrates the effectiveness of the approach, significantly improving the correctness and completeness of the generated FSMs. The study thus establishes a new paradigm for model-driven engineering that synergistically combines generative AI with formal verification techniques.
Existing FSM extraction methods suffer from limited scalability, incomplete coverage of protocol specifications, and inadequate handling of natural language ambiguity in RFC documents. To address these challenges, we propose FlowFSM—a novel framework that integrates large language model (LLM) agents, prompt chaining, and stepwise chain-of-thought reasoning to enable high-accuracy, fully automated extraction of protocol state machines from RFCs. FlowFSM decomposes the extraction task into modular subtasks, employs multi-stage rule generation, and incorporates hallucination suppression mechanisms to enhance reliability in state identification and transition inference. Experimental evaluation on FTP and RTSP protocols demonstrates a 42.7% reduction in erroneous transitions and achieves 98.3% state coverage. This work establishes a scalable, robust paradigm for protocol modeling, formal verification, and vulnerability discovery—bridging the gap between natural-language protocol specifications and precise, executable state-machine representations.
This paper addresses the automatic reverse synthesis of a finite-state machine (FSM) into an equivalent hierarchical FSM (HFSM). We propose a graph-theoretic, modular decomposition approach—introducing modular decomposition theory to automata theory for the first time—and define “thin modules” to recover algebraic structure, yielding a modular decomposition graph (MDG) that uniquely characterizes all equivalent thin HFSMs. We design an algorithm for module identification and HFSM synthesis with time complexity O(n²k), and address the bottleneck of minimizing the size of the largest component via a greedy optimization strategy. Empirical evaluation on the Harel watch benchmark demonstrates both effectiveness and scalability. Our core contribution is a formal theoretical framework for characterizing the hierarchical structure of FSMs, enabling the first reverse HFSM synthesis method that provides both rigorous formal guarantees and practical applicability.
This work addresses log-based anomaly detection in software systems. To overcome the poor interpretability and high deployment overhead of existing neural network models, we propose a behavioral modeling approach based on Probabilistic Deterministic Finite Automata (PDFA). Our method introduces an enhanced state-merging strategy that enables flexible trade-offs between model interpretability and compactness—yielding either human-readable explicit state-transition models or highly compressed, high-accuracy variants. By integrating probabilistic automaton learning, statistical hypothesis testing, and generalization optimization, our framework achieves state-of-the-art modeling accuracy across diverse real-world software logs. In anomaly detection, it attains significantly higher F1-scores than mainstream neural network baselines. Crucially, its compact model variant maintains both efficient inference latency and practical robustness, making it suitable for production deployment.
This paper addresses discrete-time interconnected systems whose subsystem dynamics and interconnection topology are partially unknown. Method: We propose a data-driven, compositional approach to construct finite-state abstractions for formal verification and distributed controller synthesis. Subsystems are modeled individually from input-output data, and—novelly—the unknown static interconnection mapping is treated as a learnable object, enabling its symbolic abstraction. Compositionality and rigorous error propagation analysis ensure that the resulting abstraction strictly satisfies an approximate simulation relation. Contribution/Results: We theoretically establish scalability and verifiability of the abstraction. Experiments demonstrate substantial mitigation of the curse of dimensionality, enabling high-precision, low-complexity controller synthesis while preserving formal guarantees.
This study addresses the inefficiency and error-proneness of manual UML state machine design, as well as the limitations of existing automated approaches in handling unstructured natural language requirements. To overcome these challenges, the authors propose a novel large language model (LLM)-based method that introduces two distinct modeling frameworks—structure-driven and event-driven—for state machine generation, complemented by a hybrid refinement strategy to iteratively optimize initial outputs. Experimental results demonstrate that Claude 3.5 Sonnet achieves F1 scores of 0.90 for states and 0.75 for transitions under a single-step prompting setup. Furthermore, the hybrid approach significantly enhances GPT-4o’s performance, bringing it close to Claude’s level, thereby validating the effectiveness and generalizability of the proposed framework.
This work addresses the challenge of reverse-engineering black-box systems that are model-free, non-resettable, and exhibit control flow dependent on internal state variables. The authors propose an active learning algorithm to infer extended finite-state machine (EFSM) models that accurately capture both data flow and control logic. Notably, this is the first method capable of effectively synthesizing EFSMs with registers and guard conditions without requiring system resets, thereby significantly relaxing assumptions about the system under test. Experimental results demonstrate that the approach achieves high-fidelity modeling of complex state-dependent systems, overcoming existing limitations in both system controllability and model expressiveness.
This work addresses the challenge of directly applying numerical time-series trajectories from cyber-physical systems (CPS) to formal verification by proposing MELA, a novel method that systematically integrates information-theoretic variable selection with decision tree–based interval abstraction to achieve fully automated, unsupervised numerical-to-symbolic conversion. By coupling this transformation with passive automata learning, MELA synthesizes compact, interpretable behavioral models from raw signals that exhibit strong correlation with underlying system states. Evaluated on two CPS case studies, MELA reduces the number of states and transitions by 49.20% on average while improving model accuracy by 41.71%, thereby effectively enabling system-level requirement verification and uncovering implicit system behaviors.