build baseline ll(1)/lr(1) parsers

Designs and implements deterministic one-token-lookahead parsers — top-down LL(1) and bottom-up LR(1) — by analyzing context-free grammars, constructing FIRST/FOLLOW sets and LR(1) item sets, and generating parse tables or corresponding parser code that performs predictive parsing or shift/reduce actions. This work includes transforming grammars (e.g., left‑factoring, eliminating left recursion), detecting and resolving conflicts (shift/reduce, reduce/reduce), integrating with a lexer, and validating correctness and error recovery on representative inputs.

buildbaselinell(1)lr(1)

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A closer look at TDFA

Jun 03, 2022
AB
A. Borsotti

This paper addresses efficient regular expression parsing and submatch extraction by proposing a unified framework based on Tagged Deterministic Finite Automata (TDFA). Methodologically, it presents for the first time a complete, systematic treatment of TDFA construction, optimization, and ambiguity resolution—including POSIX and leftmost-longest semantics—via rigorous pseudocode and step-by-step examples; supports both precompilation and on-the-fly determinization; and incorporates production-grade optimizations such as state merging and deferred transitions. Contributions include: (1) the first fully implemented TDFA solution supporting multiple semantic policies, dual determinization modes, and end-to-end optimization; and (2) empirical validation via both the RE2C generator and a standalone Java library, demonstrating substantial speedups over conventional NFA/DFA approaches across multiple benchmark suites—achieving state-of-the-art runtime performance.

Compares ahead-of-time and just-in-time determinization with performance benchmarksDevelops an algorithm for parsing and extracting submatches from regular expressionsExplains transformations from regex to optimized automaton with practical optimizations

This study addresses the longstanding trade-off between expressiveness and performance in parsing by systematically evaluating generalized context-free parsers against deterministic baselines. While deterministic parsers such as LL(1) and LR(1) constrain language design, generalized parsers offer greater expressivity but lack comprehensive empirical assessment. The authors implement six generalized algorithms—CYK, Valiant, Earley, GLL, RNGLR, and BRNGLR—in a unified Rust framework and conduct controlled benchmarks across 22 grammars ranging from arithmetic expressions to full C++ and Java specifications. Their rigorous, reproducible analysis reveals that the performance overhead of generalized parsing is substantially lower than commonly assumed: on deterministic grammars, GLR-family parsers are only about three times slower than LR(1) (median), with low variance and high stability, establishing them as the pragmatic choice for real-world applications requiring full context-free expressiveness.

deterministic parsinggeneral context-free parsinggrammar expressiveness

This work addresses the problem of syntactic ambiguity in context-free grammars, where ambiguities can lead to unintended scoping, operator precedence, and associativity that deviate from language design intent. To resolve this, the authors propose an example-based disambiguation method that synthesizes user-preferred disambiguation rules from programming examples and the original grammar using a novel tree automaton learning algorithm. They further introduce an efficient tree automaton intersection algorithm that significantly compresses the resulting specification, ensuring both readability and compatibility with mainstream parser generators. The implemented tool, Greta, successfully eliminates ambiguities across multiple case studies, producing unambiguous, canonical grammars suitable for standard parser generators, while the approach is theoretically guaranteed to be correct.

associativitycontext-free grammarsgrammar ambiguity

Constructing a BPE Tokenization DFA

May 13, 2024
MB
Martin Berglund
🏛️ Umeå University | Stellenbosch University | National Institute for Theoretical and Computational Sciences

Byte Pair Encoding (BPE) tokenization yields subword sequences that lack direct support for formal language operations, hindering rigorous pattern matching and compositional verification in open-vocabulary NLP systems. Method: We propose the first deterministic finite automaton (DFA) construction algorithm tailored to BPE output—treating tokenized sequences as constrained symbol strings without reconstructing original bytes or characters. Our approach introduces a novel equivalence-class partitioning scheme and transition function synthesis mechanism grounded in BPE merge rules, enabling linear-time O(n) DFA construction while preserving semantic fidelity. Contribution/Results: The resulting DFA supports efficient subword-level regular expression matching, lexicon equivalence checking, and formal language composition operations. It significantly improves both efficiency and composability of pattern recognition and formal verification in open-vocabulary NLP, establishing foundational automata infrastructure for verifiable, scalable, tokenization-aware language processing.

Analyzing state complexity of tokenization automataEfficient DFA construction for BPE tokenizationEnabling pattern matching on tokenized text

Latest Papers

What's happening recently
View more

This work addresses the inefficiency of traditional backtracking-based regular expression engines, which suffer from poor worst-case performance and struggle to support advanced features such as lookaheads, negative lookaheads, and submatch extraction. Focusing on regular expressions with lookahead and negative lookahead (REwLA), the paper presents the first finite automaton transformation framework capable of handling submatches. By leveraging formal language theory and automata construction techniques, it converts an REwLA of size \( m \) into a deterministic finite automaton (DFA) with \( \tilde{O}(2^{2^m}) \) states, and further extends this approach to weighted REwLA by constructing a weighted NFA with comparable state complexity. This method overcomes the performance limitations of backtracking and establishes both a theoretical foundation and a practical pathway for efficiently implementing advanced regular expression functionalities.

finite automatonlookaheadmatching algorithm

This work addresses the challenge that large language models often generate code violating the syntax of domain-specific languages (DSLs) when invoking external services, a problem exacerbated by the absence of context-free grammars for third-party DSLs needed for syntax-constrained decoding. To overcome this, the authors propose Autogrammar, an agent that uniquely integrates Kripke structures with language models, enabling declarative control of agent behavior via linear temporal logic and automatically inducing DSL grammars from documentation and execution feedback—eliminating the need for manual grammar engineering. Evaluated on three real-world DSLs, the learned grammars achieve near-perfect precision (≈100%) on unseen data, significantly outperforming baseline methods in end-to-end task accuracy, matching or exceeding human-crafted grammars in 80% of tasks, while accelerating inference by 3.8×.

context-free grammardomain-specific languagegrammar-constrained decoding

This work addresses the distortion of a language model’s original probability distribution caused by rigid masking in grammar-constrained decoding, which often yields valid but suboptimal outputs, while existing distribution-recovery methods incur substantial computational overhead. The authors propose a lightweight, offline-trained logit correction approach that leverages lexical and parser internal states—along with candidate next tokens—from the incremental parsing process as priors to recover the true distribution without modifying model weights. This method achieves effective correction at zero additional inference cost and implicitly captures lookahead effects using only candidate tokens. Experiments across multiple grammars demonstrate that the proposed technique significantly reduces the divergence between masked and true distributions, consistently outperforming both pure masking and online resampling baselines, with even its lightest variant matching or surpassing their performance.

bias correctionconstrained generationGrammar Constrained Decoding

This work addresses the high latency of large language models in context-free grammar (CFG)-constrained decoding, which stems from the need to traverse the full vocabulary at each generation step, rendering complex grammars impractical. To overcome this limitation, the authors propose CFGzip, a novel offline method that compresses the token search space by integrating CFG analysis with semantic token clustering to construct a compact yet complete subset of valid tokens. This approach preserves generation correctness while drastically reducing the search scope. CFGzip seamlessly integrates with existing grammar-constrained decoding engines and, in standard evaluations, achieves up to a 7.5× speedup and reduces latency by two orders of magnitude, substantially enhancing the practicality and scalability of generating text under complex CFG constraints.

constrained decodingcontext-free grammardecoding overhead