parsing

Designs, implements, and evaluates algorithms, grammars, and software that analyze token streams or text to produce structured representations such as parse trees, abstract syntax trees, or dependency graphs; this includes building lexers, grammar specifications, parser generators or hand-written parsers, and mechanisms for error recovery, ambiguity resolution, incremental/streaming parsing, performance and robustness. It also covers transforming parse structures into downstream representations, testing and debugging grammars, and integrating parsing into larger processing pipelines.

parsing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.78
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$183K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Automating the Analysis of Parsing Algorithms (and other Dynamic Programs)

Dec 29, 2025
TV
Tim Vieira
🏛️ Johns Hopkins University | ETH Zürich

This paper addresses the challenge of establishing performance guarantees for dynamic programming (DP) parsing algorithms in natural language processing. We present the first automated analysis system that unifies program analysis and complexity inference within a DP framework. Our approach integrates static analysis, type inference, abstract interpretation, and dependency graph modeling to enable formal verification and synthesis of efficient data structures. Key contributions include: (1) a unified formal model capturing DP control flow, data flow, and recurrence structure; (2) automatic inference of precise types, detection of dead code, and identification of redundant computations; and (3) generation of tight, parameterized upper bounds on time and space complexity. We evaluate our system on canonical parsing algorithms—including CKY, Earley, and Neural PCFG—demonstrating substantial improvements in both the automation level and precision of complexity analysis.

Automating analysis of parsing algorithms and dynamic programsInferring types, dead code, and verifying algorithm propertiesProviding guarantees on runtime and space complexity bounds

Inferring Attributed Grammars from Parser Implementations

Jul 17, 2025
AP
Andreas Pointner
🏛️ University of Applied Sciences, Upper Austria | Johannes Kepler University Linz

Existing structured-input processing systems often lack complete and up-to-date syntactic and semantic specifications; while syntax mining has focused primarily on parsing structure, semantic recovery remains unaddressed. Method: We propose the first approach to automatically infer attribute grammars from recursive-descent parser implementations. Our method combines dynamic execution tracing and program instrumentation to capture runtime behavior, augmented by control-flow analysis and grammar-driven semantic mapping, thereby precisely associating parsing operations with productions and extracting semantic actions. Contribution/Results: This work pioneers syntax mining at the semantic level, enabling fully automated generation of executable attribute grammars that faithfully model input-processing logic. Evaluation across multiple real-world programs demonstrates that the inferred grammars accurately reproduce original parser behavior—enabling novel applications in reverse engineering, specification documentation, and security analysis.

Inferring attributed grammars from parser implementationsMapping runtime behavior to grammar for specification recoveryRecovering semantic aspects of input handling from parsers

Generating Inputs for Grammar Mining using Dynamic Symbolic Execution

Aug 05, 2025
AP
Andreas Pointner
🏛️ University of Applied Sciences Upper Austria | Johannes Kepler University

In grammar reverse-engineering of legacy parsers, insufficient input samples often lead to incomplete grammar coverage—particularly missing edge cases or deprecated features. Method: This paper proposes an automated input generation approach based on dynamic symbolic execution (DSE), the first to apply DSE to grammar mining. We design a three-stage decoupled input generation framework and an iterative expansion strategy to effectively mitigate DSE’s inherent limitations in handling structured inputs. Crucially, our method requires no prior input samples and systematically triggers deep parser behaviors. Results: Evaluated on 11 real-world benchmarks, our generated grammars achieve precision and recall comparable to state-of-the-art methods, while significantly improving detection of subtle semantic features and historical edge-case usages.

Generating diverse inputs for grammar mining automaticallyImproving grammar coverage by capturing edge casesOvercoming limitations of Dynamic Symbolic Execution for parsers

A Systematic Comparison of Syntactic Representations of Dependency Parsing

May 29, 2017
GW
Guillaume Wisniewski
🏛️ Univ. Paris-Sud | Université Paris-Saclay | University of Copenhagen

This study systematically investigates how dependency annotation schemes affect the performance of transition-based parsers. Method: Addressing language-specific non-canonical structures in Universal Dependencies (UD) treebanks, we design standardization transformation rules and comparatively evaluate parser performance—measured by LAS and UAS—under both original and standardized annotations within a unified, multilingual evaluation framework. Contribution/Results: We empirically demonstrate, for the first time, that annotation standardization does not universally improve parsing accuracy. Crucially, we reveal that linguistic typological features significantly moderate the effectiveness of annotation schemes: for certain languages, the original non-standard annotations yield higher accuracy than standardized ones. This finding challenges the implicit assumption that standardization is inherently optimal and underscores the necessity of considering language-specific syntactic properties when selecting or designing syntactic representations.

Compare parser performance across annotation schemes.Convert syntactic constructions to standard representations.Evaluate parsing performance across multiple languages.

XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models

Nov 22, 2024
YD
Yixin Dong
🏛️ Carnegie Mellon University | NVIDIA | Shanghai Jiao Tong University | University of California, Berkeley

Large language model (LLM) agents increasingly require structured, parseable outputs—such as code, function calls, or embodied instructions—yet conventional context-free grammar (CFG)-based constrained decoding suffers from high computational overhead due to full-vocabulary traversal and multi-stack state management. Method: This paper introduces a CFG-driven constrained decoding engine that is context-free grammar–aware and computationally efficient. We propose a novel “lexical divide-and-conquer” strategy: pre-filtering context-free tokens while dynamically interpreting context-sensitive ones; further, we design grammar-context expansion transformations and a persistent stack mechanism to tightly couple grammar computation with GPU-based inference in a pipelined fashion. Contribution/Results: Experiments demonstrate up to 100× speedup over state-of-the-art approaches. The engine achieves near-zero-overhead structured generation in end-to-end low-latency serving, significantly improving the reliability and real-time responsiveness of LLM agents on complex, structured tasks.

Accelerates grammar computation with GPU co-designEnables efficient structured output generation for LLMsReduces overhead in context-free grammar execution

Latest Papers

What's happening recently
View more

This work addresses the problem of syntactic ambiguity in context-free grammars, where ambiguities can lead to unintended scoping, operator precedence, and associativity that deviate from language design intent. To resolve this, the authors propose an example-based disambiguation method that synthesizes user-preferred disambiguation rules from programming examples and the original grammar using a novel tree automaton learning algorithm. They further introduce an efficient tree automaton intersection algorithm that significantly compresses the resulting specification, ensuring both readability and compatibility with mainstream parser generators. The implemented tool, Greta, successfully eliminates ambiguities across multiple case studies, producing unambiguous, canonical grammars suitable for standard parser generators, while the approach is theoretically guaranteed to be correct.

associativitycontext-free grammarsgrammar ambiguity

This study addresses the longstanding trade-off between expressiveness and performance in parsing by systematically evaluating generalized context-free parsers against deterministic baselines. While deterministic parsers such as LL(1) and LR(1) constrain language design, generalized parsers offer greater expressivity but lack comprehensive empirical assessment. The authors implement six generalized algorithms—CYK, Valiant, Earley, GLL, RNGLR, and BRNGLR—in a unified Rust framework and conduct controlled benchmarks across 22 grammars ranging from arithmetic expressions to full C++ and Java specifications. Their rigorous, reproducible analysis reveals that the performance overhead of generalized parsing is substantially lower than commonly assumed: on deterministic grammars, GLR-family parsers are only about three times slower than LR(1) (median), with low variance and high stability, establishing them as the pragmatic choice for real-world applications requiring full context-free expressiveness.

deterministic parsinggeneral context-free parsinggrammar expressiveness

This work addresses the challenge of identifying structural and semantic similarities across imperative programs written in different languages by proposing a unified graph representation that integrates abstract syntax trees with neural semantic embeddings. The approach transforms annotated programs into typed, attributed graphs and leverages CodeBERT and SentenceTransformer to generate rich semantic embeddings. By constructing consistent graph representations across multilingual verification datasets—including C/ACSL, Java/JML, and Dafny—it achieves, for the first time, joint modeling of syntactic structure and formal semantics. This unified framework offers a viable pathway for cross-language reuse of verification artifacts and demonstrates strong generality and effectiveness across diverse programming languages and specification frameworks.

graph constructionimperative programsprogram representation

This study investigates the efficient development of formal grammatical resources for Cantonese and Irish while preserving cross-linguistic consistency, and evaluates the potential of multilingual large language models (LLMs) to assist in grammar engineering for low-resource languages. Building upon the ParGram framework, we present the first systematic application of multilingual LLMs—such as gpt-oss-120b—to parallel treebank construction, leveraging model-assisted translation and syntactic structure generation to maintain alignment at the level of abstract functional representations. Experimental results indicate that model-generated translations show limited efficacy and are unaffected by the choice of prompt language; although syntactic generation captures predicate–argument relations to some extent, it underperforms on cross-linguistically abstract tasks, necessitating expert intervention. This work establishes a novel paradigm and empirical foundation for formal syntactic modeling of low-resource languages.

Cantonesegrammar engineeringIrish

This study addresses the problem of context-free grammar (CFG)-constrained reachability queries over graphs. To this end, the authors propose an algorithmic framework that balances theoretical efficiency with practical performance, featuring sub-cubic time preprocessing, multiple indexing strategies, and a structured decoding semantics. Through systematic experiments on real-world graph datasets, the effectiveness of the approach is empirically validated. The study further elucidates how the structure of the grammar and the characteristics of the underlying graph jointly influence query performance, and quantifies the trade-offs between index construction overhead and query efficiency. These findings offer both theoretical insights and practical guidance for selecting appropriate algorithms in real-world applications involving CFG-constrained graph reachability.

context-free grammarcontext-free language reachabilitygrammar-constrained reachability