Score
Design, implement, and analyze evaluation methods, metrics, and tools that assess the fidelity and quality of interaction finite-state machines (FSMs). This includes comparing generated FSMs to reference/ideal FSMs using graph-structural measures, embedding-based semantic similarity, and producing quantitative interaction-quality or interactivity scores.
This work addresses the challenge of generating reliable system models directly from natural language requirements by proposing a novel framework that leverages large language models (LLMs) for the automatic generation and repair of finite state machines (FSMs). It introduces, for the first time, the use of LLMs in FSM construction and innovatively integrates FSM mutation analysis with automated test generation to establish an expert feedback–driven optimization mechanism. Experimental validation on synthetic datasets using GPT-4 demonstrates the effectiveness of the approach, significantly improving the correctness and completeness of the generated FSMs. The study thus establishes a new paradigm for model-driven engineering that synergistically combines generative AI with formal verification techniques.
Students commonly struggle to grasp the operational semantics of nondeterministic finite automata (NFAs) and pushdown automata (PDAs), particularly regarding multi-path computation, stack-state dependencies, and distinguishing configurations with identical control states but differing stack contents. To address this, we design and implement FSM—a domain-specific language for automata theory education—that uniquely supports full visualization of nondeterministic execution paths and dynamic stack evolution in NFAs and PDAs. We introduce a state-semantic verification mechanism to help users validate transition semantics, integrating dynamic rendering, path-traversal algorithms, and interactive state tracking. Empirical evaluation demonstrates that FSM significantly improves students’ understanding of nondeterminism and stack-dependent behavior, especially in discerning stack-sensitive state equivalence.
Existing approaches struggle to effectively evaluate the quality of dynamic interactions in AI-generated explorable explanations, particularly lacking metrics for learner-controlled state transitions and context-sensitive feedback. This work proposes EE-Eval, a novel framework that formalizes interactivity as a finite state machine (FSM) and enables automated, multidimensional assessment by comparing the structural and semantic alignment between generated content and an ideal pedagogical FSM grounded in instructional intent. Integrating graph similarity with embedding-based semantic analysis, EE-Eval overcomes the limitations of conventional methods that focus narrowly on code executability or visual fidelity. Extensive experiments across 127 concepts and thousands of explanations generated by six AI models demonstrate that EE-Eval significantly outperforms existing baselines and exhibits strong agreement with human judgments of both interactivity and instructional effectiveness.
Existing approaches struggle to model the latent states that drive interactions in role-playing scenarios, often resulting in inconsistent behaviors. This work proposes a novel method that integrates large language models with finite state machines to automatically generate interpretable and executable deterministic finite state machines (CFSMs) from unstructured character descriptions. The approach is further extended to probabilistic finite state machines (CPFSMs) to capture character states and their transitions more flexibly. By enabling structured yet open-domain modeling of character behavior, the method significantly outperforms general-purpose baselines in both synthetic evaluations and real-world role-playing settings, demonstrating its effectiveness in handling structured tasks and stochastic state exploration.
Traditional binary correctness verification fails to capture quantitative system behaviors. Method: We propose the first automated toolkit for quantitative automata supporting six classical semantics—Inf, Sup, LimInf, LimSup, LimInfAvg, and LimSupAvg—and systematically address core decision problems: emptiness, inclusion, equivalence, and safety/liveness verification. Our approach introduces weighted transition modeling and a generalized value-function framework, integrating symbolic decision procedures, optimization solvers, and automata transformation techniques to enable extremal-value computation, safety-liveness decomposition, and real-time monitoring. Contribution/Results: Experiments demonstrate efficiency on inclusion checking, constant-function recognition, and online monitoring tasks. We release the first open-source benchmark suite for quantitative automata analysis, establishing a scalable, modular, and unified infrastructure for quantitative system verification.
本文通过设计4个任务和1482个查询来评估大语言模型对网络协议状态机的理解能力,研究其在形式化表示中的准确性与影响因素。
This study addresses the inefficiency and error-proneness of manual UML state machine design, as well as the limitations of existing automated approaches in handling unstructured natural language requirements. To overcome these challenges, the authors propose a novel large language model (LLM)-based method that introduces two distinct modeling frameworks—structure-driven and event-driven—for state machine generation, complemented by a hybrid refinement strategy to iteratively optimize initial outputs. Experimental results demonstrate that Claude 3.5 Sonnet achieves F1 scores of 0.90 for states and 0.75 for transitions under a single-step prompting setup. Furthermore, the hybrid approach significantly enhances GPT-4o’s performance, bringing it close to Claude’s level, thereby validating the effectiveness and generalizability of the proposed framework.
Existing benchmarks for interactive agents struggle to simultaneously ensure scalability and effectively evaluate performance under realistic workflow conditions involving state conflicts—such as partial, stale, or contradictory prior states. To address this gap, this work proposes ClawForge, the first executable command-line workflow framework that systematically supports evaluation under pre-existing state conflicts. ClawForge enables reproducible task construction through scenario templates, state initialization, reference trajectories, and validators, while abandoning strict trajectory matching in favor of stepwise assessment based on normalized final states and observable side effects. The accompanying ClawForge-Bench comprises 17 scenarios; evaluations reveal that even the best-performing model achieves only a 45.3% strict accuracy, with all models exhibiting errorful state replacement rates below 17%. Critically, the tendency to proactively inspect existing states emerges as a key determinant of performance disparities.
This work investigates the capability of large language models (LLMs) to recover finite-state machine (FSM) behavior from natural language descriptions and generate correct register-transfer level (RTL) code. To this end, we introduce the first fully automated and scalable FSM-to-RTL benchmark, comprising over a thousand test cases, with data quality ensured through a structured YAML intermediate representation, formal verification via SAT solvers, and manual validation. Our experiments demonstrate that supervised fine-tuning substantially improves out-of-distribution generalization, while test-time compute scaling enhances reasoning reliability. Nevertheless, even the most advanced LLMs exhibit a significant drop in accuracy on complex FSM tasks, revealing fundamental limitations in current models’ ability to reason about hardware semantics.
This study addresses the unclear relationships among information sources, generation strategies, and quality evidence in test case generation using large language models (LLMs). Through a systematic literature review of 95 studies, this work constructs a multidimensional taxonomy and a benchmark analysis framework. Specifically, it proposes a four-dimensional classification system that elucidates how execution feedback influences oracle independence. Furthermore, it establishes a unified theoretical framework connecting the generation process with quality assessment. By identifying independent oracle evaluation as a critical yet underexplored dimension, this research formulates a future agenda centered on rigorous, oracle-independent quality measurement for LLM-generated test cases.