Score
Designs, implements, and evaluates algorithms, grammars, and software that analyze token streams or text to produce structured representations such as parse trees, abstract syntax trees, or dependency graphs; this includes building lexers, grammar specifications, parser generators or hand-written parsers, and mechanisms for error recovery, ambiguity resolution, incremental/streaming parsing, performance and robustness. It also covers transforming parse structures into downstream representations, testing and debugging grammars, and integrating parsing into larger processing pipelines.
This paper addresses the challenge of establishing performance guarantees for dynamic programming (DP) parsing algorithms in natural language processing. We present the first automated analysis system that unifies program analysis and complexity inference within a DP framework. Our approach integrates static analysis, type inference, abstract interpretation, and dependency graph modeling to enable formal verification and synthesis of efficient data structures. Key contributions include: (1) a unified formal model capturing DP control flow, data flow, and recurrence structure; (2) automatic inference of precise types, detection of dead code, and identification of redundant computations; and (3) generation of tight, parameterized upper bounds on time and space complexity. We evaluate our system on canonical parsing algorithms—including CKY, Earley, and Neural PCFG—demonstrating substantial improvements in both the automation level and precision of complexity analysis.
Existing structured-input processing systems often lack complete and up-to-date syntactic and semantic specifications; while syntax mining has focused primarily on parsing structure, semantic recovery remains unaddressed. Method: We propose the first approach to automatically infer attribute grammars from recursive-descent parser implementations. Our method combines dynamic execution tracing and program instrumentation to capture runtime behavior, augmented by control-flow analysis and grammar-driven semantic mapping, thereby precisely associating parsing operations with productions and extracting semantic actions. Contribution/Results: This work pioneers syntax mining at the semantic level, enabling fully automated generation of executable attribute grammars that faithfully model input-processing logic. Evaluation across multiple real-world programs demonstrates that the inferred grammars accurately reproduce original parser behavior—enabling novel applications in reverse engineering, specification documentation, and security analysis.
In grammar reverse-engineering of legacy parsers, insufficient input samples often lead to incomplete grammar coverage—particularly missing edge cases or deprecated features. Method: This paper proposes an automated input generation approach based on dynamic symbolic execution (DSE), the first to apply DSE to grammar mining. We design a three-stage decoupled input generation framework and an iterative expansion strategy to effectively mitigate DSE’s inherent limitations in handling structured inputs. Crucially, our method requires no prior input samples and systematically triggers deep parser behaviors. Results: Evaluated on 11 real-world benchmarks, our generated grammars achieve precision and recall comparable to state-of-the-art methods, while significantly improving detection of subtle semantic features and historical edge-case usages.
This study systematically investigates how dependency annotation schemes affect the performance of transition-based parsers. Method: Addressing language-specific non-canonical structures in Universal Dependencies (UD) treebanks, we design standardization transformation rules and comparatively evaluate parser performance—measured by LAS and UAS—under both original and standardized annotations within a unified, multilingual evaluation framework. Contribution/Results: We empirically demonstrate, for the first time, that annotation standardization does not universally improve parsing accuracy. Crucially, we reveal that linguistic typological features significantly moderate the effectiveness of annotation schemes: for certain languages, the original non-standard annotations yield higher accuracy than standardized ones. This finding challenges the implicit assumption that standardization is inherently optimal and underscores the necessity of considering language-specific syntactic properties when selecting or designing syntactic representations.
Large language model (LLM) agents increasingly require structured, parseable outputs—such as code, function calls, or embodied instructions—yet conventional context-free grammar (CFG)-based constrained decoding suffers from high computational overhead due to full-vocabulary traversal and multi-stack state management. Method: This paper introduces a CFG-driven constrained decoding engine that is context-free grammar–aware and computationally efficient. We propose a novel “lexical divide-and-conquer” strategy: pre-filtering context-free tokens while dynamically interpreting context-sensitive ones; further, we design grammar-context expansion transformations and a persistent stack mechanism to tightly couple grammar computation with GPU-based inference in a pipelined fashion. Contribution/Results: Experiments demonstrate up to 100× speedup over state-of-the-art approaches. The engine achieves near-zero-overhead structured generation in end-to-end low-latency serving, significantly improving the reliability and real-time responsiveness of LLM agents on complex, structured tasks.
This work addresses the problem of syntactic ambiguity in context-free grammars, where ambiguities can lead to unintended scoping, operator precedence, and associativity that deviate from language design intent. To resolve this, the authors propose an example-based disambiguation method that synthesizes user-preferred disambiguation rules from programming examples and the original grammar using a novel tree automaton learning algorithm. They further introduce an efficient tree automaton intersection algorithm that significantly compresses the resulting specification, ensuring both readability and compatibility with mainstream parser generators. The implemented tool, Greta, successfully eliminates ambiguities across multiple case studies, producing unambiguous, canonical grammars suitable for standard parser generators, while the approach is theoretically guaranteed to be correct.
This study addresses the longstanding trade-off between expressiveness and performance in parsing by systematically evaluating generalized context-free parsers against deterministic baselines. While deterministic parsers such as LL(1) and LR(1) constrain language design, generalized parsers offer greater expressivity but lack comprehensive empirical assessment. The authors implement six generalized algorithms—CYK, Valiant, Earley, GLL, RNGLR, and BRNGLR—in a unified Rust framework and conduct controlled benchmarks across 22 grammars ranging from arithmetic expressions to full C++ and Java specifications. Their rigorous, reproducible analysis reveals that the performance overhead of generalized parsing is substantially lower than commonly assumed: on deterministic grammars, GLR-family parsers are only about three times slower than LR(1) (median), with low variance and high stability, establishing them as the pragmatic choice for real-world applications requiring full context-free expressiveness.
This work addresses the challenge of identifying structural and semantic similarities across imperative programs written in different languages by proposing a unified graph representation that integrates abstract syntax trees with neural semantic embeddings. The approach transforms annotated programs into typed, attributed graphs and leverages CodeBERT and SentenceTransformer to generate rich semantic embeddings. By constructing consistent graph representations across multilingual verification datasets—including C/ACSL, Java/JML, and Dafny—it achieves, for the first time, joint modeling of syntactic structure and formal semantics. This unified framework offers a viable pathway for cross-language reuse of verification artifacts and demonstrates strong generality and effectiveness across diverse programming languages and specification frameworks.
This study investigates the efficient development of formal grammatical resources for Cantonese and Irish while preserving cross-linguistic consistency, and evaluates the potential of multilingual large language models (LLMs) to assist in grammar engineering for low-resource languages. Building upon the ParGram framework, we present the first systematic application of multilingual LLMs—such as gpt-oss-120b—to parallel treebank construction, leveraging model-assisted translation and syntactic structure generation to maintain alignment at the level of abstract functional representations. Experimental results indicate that model-generated translations show limited efficacy and are unaffected by the choice of prompt language; although syntactic generation captures predicate–argument relations to some extent, it underperforms on cross-linguistically abstract tasks, necessitating expert intervention. This work establishes a novel paradigm and empirical foundation for formal syntactic modeling of low-resource languages.
This study addresses the problem of context-free grammar (CFG)-constrained reachability queries over graphs. To this end, the authors propose an algorithmic framework that balances theoretical efficiency with practical performance, featuring sub-cubic time preprocessing, multiple indexing strategies, and a structured decoding semantics. Through systematic experiments on real-world graph datasets, the effectiveness of the approach is empirically validated. The study further elucidates how the structure of the grammar and the characteristics of the underlying graph jointly influence query performance, and quantifies the trade-offs between index construction overhead and query efficiency. These findings offer both theoretical insights and practical guidance for selecting appropriate algorithms in real-world applications involving CFG-constrained graph reachability.