Score
Designs, implements, and evaluates parsers and parser infrastructure driven by formal specifications, translating spec fields into grammar or regex rules, implementing syntax and semantic checks, and producing incremental, pattern-based, or trace-oriented parsing components. Builds unified or shared parser suites and reference implementations, generates parser tests from specifications, and constructs comparison frameworks to validate and verify parser outputs against constraints.
In grammar reverse-engineering of legacy parsers, insufficient input samples often lead to incomplete grammar coverage—particularly missing edge cases or deprecated features. Method: This paper proposes an automated input generation approach based on dynamic symbolic execution (DSE), the first to apply DSE to grammar mining. We design a three-stage decoupled input generation framework and an iterative expansion strategy to effectively mitigate DSE’s inherent limitations in handling structured inputs. Crucially, our method requires no prior input samples and systematically triggers deep parser behaviors. Results: Evaluated on 11 real-world benchmarks, our generated grammars achieve precision and recall comparable to state-of-the-art methods, while significantly improving detection of subtle semantic features and historical edge-case usages.
Existing structured-input processing systems often lack complete and up-to-date syntactic and semantic specifications; while syntax mining has focused primarily on parsing structure, semantic recovery remains unaddressed. Method: We propose the first approach to automatically infer attribute grammars from recursive-descent parser implementations. Our method combines dynamic execution tracing and program instrumentation to capture runtime behavior, augmented by control-flow analysis and grammar-driven semantic mapping, thereby precisely associating parsing operations with productions and extracting semantic actions. Contribution/Results: This work pioneers syntax mining at the semantic level, enabling fully automated generation of executable attribute grammars that faithfully model input-processing logic. Evaluation across multiple real-world programs demonstrates that the inferred grammars accurately reproduce original parser behavior—enabling novel applications in reverse engineering, specification documentation, and security analysis.
Formal program specifications are notoriously difficult, error-prone, and inefficient to write manually. To address this, we propose a two-stage LLM-driven approach: dialogue-guided specification synthesis followed by mutation-based verification. First, multi-turn dialogues model complex semantic requirements; second, four mutation operators—insertion, replacement, deletion, and reordering—enable verifiability-driven selection, eliminating reliance on rigid templates or syntactic grammars. Our method integrates code understanding, prompt engineering, and heuristic verifiability assessment. Evaluated on SV-COMP and a custom Java benchmark comprising 385 programs, it generates 279 verifiable specifications. These achieve significantly higher completeness and accuracy than pure-LLM baselines and classical tools (e.g., Houdini, Daikon). To our knowledge, this is the first approach to achieve both high coverage and formal verifiability in fully automated specification generation.
Existing grammar inference techniques fail to accurately reverse-engineer context-free grammars from real-world recursive-descent parsers lacking formal syntactic specifications. Method: We propose the first symbolic grammar mining approach tailored to industrial-grade recursive-descent parsers. Our method statically models parser semantics via program analysis and integrates nonterminal mapping, path truncation, and bounded recursion unfolding to mitigate path explosion and infinite recursion—enabling fully automated, seedless inference of precise context-free grammars. Contribution/Results: Evaluated on complex parsers including TINY-C and JSON, our technique achieves 99–100% grammar extraction accuracy, substantially outperforming prior work. This is the first end-to-end grammar reverse-engineering solution for production recursive-descent parsers, enabling grammar-based full-coverage test generation, protocol reverse engineering, and automatic documentation.
Addressing the “oracle absence” and “error attribution difficulty” challenges in network protocol parser verification, this paper proposes an LLM-driven framework for RFC semantic parsing and feedback-based oracle refinement. First, large language models automatically translate unstructured RFC text into formal message specifications. Second, an iterative, quasi-oracle is constructed to support specification-guided fuzz testing and cross-language (C/Python/Go) protocol implementation verification. Finally, vulnerabilities are precisely traced back to their originating RFC clauses. This work is the first to integrate LLM-based semantic understanding with dynamic oracle refinement. Evaluated on nine mainstream protocols, it discovers 69 vulnerabilities—36 of which have been confirmed—surpassing state-of-the-art approaches in both effectiveness and efficiency. It also demonstrates, for the first time, the feasibility of fully automated derivation of test oracles directly from natural-language protocol specifications.
This study addresses the longstanding trade-off between expressiveness and performance in parsing by systematically evaluating generalized context-free parsers against deterministic baselines. While deterministic parsers such as LL(1) and LR(1) constrain language design, generalized parsers offer greater expressivity but lack comprehensive empirical assessment. The authors implement six generalized algorithms—CYK, Valiant, Earley, GLL, RNGLR, and BRNGLR—in a unified Rust framework and conduct controlled benchmarks across 22 grammars ranging from arithmetic expressions to full C++ and Java specifications. Their rigorous, reproducible analysis reveals that the performance overhead of generalized parsing is substantially lower than commonly assumed: on deterministic grammars, GLR-family parsers are only about three times slower than LR(1) (median), with low variance and high stability, establishing them as the pragmatic choice for real-world applications requiring full context-free expressiveness.
This work addresses the challenge of ensuring program correctness in natural language-to-code generation, which is often hindered by the absence of high-quality formal specifications. The authors propose VeriSpecGen, a framework that decomposes natural language requirements into atomic clauses through a traceable refinement mechanism, generates requirement-driven tests with explicit traceability mappings, and synthesizes formal specifications aligned with user intent by localizing and repairing faulty clauses upon verification failure. Integrating large language models (e.g., Claude Opus 4.5) with the Lean proof assistant, the approach leverages refinement trajectories to generate 343K training samples, substantially enhancing model generalization and reasoning capabilities. Evaluated on the VERINA SpecGen benchmark, VeriSpecGen achieves an accuracy of 86.6%, outperforming the best baseline by up to 31.8 percentage points and demonstrating a relative improvement of 62–106% in specification synthesis performance.
Existing testing approaches for regular expression engines are hindered by syntactic discrepancies across dialects and inefficient fuzzing strategies, limiting their ability to uncover defects effectively. This work proposes ReTest, a novel framework that integrates grammar-aware fuzzing with 16 Kleene algebra–based homomorphic metamorphic relations, enabling the first systematic, cross-dialect, dependency-free testing methodology. By modeling regex semantics and employing coverage-guided exploration, ReTest achieves threefold higher edge coverage than state-of-the-art techniques on PCRE and discovers three previously unknown memory-safety vulnerabilities, substantially advancing the reliability validation of regex engines.
This work addresses the high cost of manually writing formal specifications and the limitations of existing large language model (LLM)-based approaches that require white-box access to source code, thereby posing intellectual property and deployment constraints. The authors propose a black-box-driven method that leverages only test code and dynamic execution traces to generate candidate Java Modeling Language (JML) specifications via an LLM. These candidates are locally validated using bounded model checking, and an iterative feedback loop refines them based on verification outcomes. This approach is the first to enable fully automated formal specification generation without any access to the program’s internal structure. Evaluated on the SpecGenBench benchmark, it demonstrates that test-derived information effectively guides specification synthesis, while also highlighting critical challenges in checker compatibility and diagnostic feedback, substantially enhancing industrial applicability.