automaton-constrained generation

Designs, builds, or analyzes sequence generators that are constrained by deterministic finite automata (DFAs): construct DFA encodings of sequence constraints and implement generation mechanisms that traverse DFA states to produce only valid output words or event sequences (typically deterministically).

automaton-constrainedgeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the problem of automatically generating event sequences with specified dependencies. To this end, the authors propose a formal method that integrates deterministic finite automata (DFA) with finite-state transducers, presenting the first systematic framework that combines these two formalisms for synthesizing dependency-aware event sequences driven by multiple inputs. The approach is implemented in AGDES, a configurable tool capable of automatically producing output sequences that adhere to given dependency constraints. Empirical evaluation demonstrates that AGDES exhibits strong applicability and practical utility in scenarios involving formal verification and testing, offering a robust solution for generating semantically coherent and constraint-compliant event traces.

Automatic GenerationDependent Event SequencesDeterministic Finite Automaton

Constructing a BPE Tokenization DFA

May 13, 2024
MB
Martin Berglund
🏛️ Umeå University | Stellenbosch University | National Institute for Theoretical and Computational Sciences

Byte Pair Encoding (BPE) tokenization yields subword sequences that lack direct support for formal language operations, hindering rigorous pattern matching and compositional verification in open-vocabulary NLP systems. Method: We propose the first deterministic finite automaton (DFA) construction algorithm tailored to BPE output—treating tokenized sequences as constrained symbol strings without reconstructing original bytes or characters. Our approach introduces a novel equivalence-class partitioning scheme and transition function synthesis mechanism grounded in BPE merge rules, enabling linear-time O(n) DFA construction while preserving semantic fidelity. Contribution/Results: The resulting DFA supports efficient subword-level regular expression matching, lexicon equivalence checking, and formal language composition operations. It significantly improves both efficiency and composability of pattern recognition and formal verification in open-vocabulary NLP, establishing foundational automata infrastructure for verifiable, scalable, tokenization-aware language processing.

Analyzing state complexity of tokenization automataEfficient DFA construction for BPE tokenizationEnabling pattern matching on tokenized text

Traditional deterministic finite automata (DFA) employ uniform, unweighted state transitions, limiting their capacity to model quantitative or graded behaviors. Method: This paper introduces the *discharging deterministic finite automaton* (DDFA), the first automaton model incorporating graph-theoretic discharging principles: states “discharge” rational-valued weights to adjacent states upon reading input symbols, enabling dynamic, weighted transitions. We formally define DDFA and analyze its algebraic structure. Contribution/Results: We prove that the set of sequences generated by a DDFA forms a ring—the *quasi-k-regular sequence ring*—which strictly generalizes the classical k-regular sequence ring of Allouche and Shallit. Crucially, quasi-k-regularity properly contains k-regularity and exhibits superior algebraic completeness. By integrating formal language theory, directed graph modeling, ring theory, and discrete dynamical systems, this work significantly extends the algebraic expressiveness and applicability of automata models.

Extends k-regular sequences to quasi-k-regular ringGeneralizes DFAs using graph theory's discharging methodIntroduces Discharging Deterministic Finite Automata (DDFA) structure

A Linear-time Simulation of Deterministic d-Limited Automata

Dec 04, 2023
AR
Alexander Rubtsov
🏛️ National Research University Higher School of Economics

This paper investigates language recognition for deterministic $d$-finite automata (DLAs) and their generalized $d(n)$-finite automata. For membership testing, we present the first strictly linear-time $O(n)$ recognition algorithm under the RAM model, overcoming a long-standing efficiency gap in the Hibbard–Pighizzini hierarchy. Our method integrates state compression encoding, input-driven traversal scheduling, and dynamic access-count management, supporting arbitrary computable memory-access bound functions $d(n)$. Unlike prior approaches limited to constant $d$, our algorithm uniformly and efficiently simulates both constant and non-constant $d(n)$, enabling hardware-friendly parsing for subclasses of deterministic context-free languages (DCFLs), such as LR(1) languages. This bridges the theoretical–practical divide between abstract automaton models and real-world linear-time parsing, establishing a novel, implementable paradigm grounded in rigorous automata theory.

Extending linear-time recognition beyond deterministic context-free languagesLinear-time recognition algorithm for deterministic d-limited automataSolving membership problem for deterministic d(n)-limited automata

A closer look at TDFA

Jun 03, 2022
AB
A. Borsotti

This paper addresses efficient regular expression parsing and submatch extraction by proposing a unified framework based on Tagged Deterministic Finite Automata (TDFA). Methodologically, it presents for the first time a complete, systematic treatment of TDFA construction, optimization, and ambiguity resolution—including POSIX and leftmost-longest semantics—via rigorous pseudocode and step-by-step examples; supports both precompilation and on-the-fly determinization; and incorporates production-grade optimizations such as state merging and deferred transitions. Contributions include: (1) the first fully implemented TDFA solution supporting multiple semantic policies, dual determinization modes, and end-to-end optimization; and (2) empirical validation via both the RE2C generator and a standalone Java library, demonstrating substantial speedups over conventional NFA/DFA approaches across multiple benchmark suites—achieving state-of-the-art runtime performance.

Compares ahead-of-time and just-in-time determinization with performance benchmarksDevelops an algorithm for parsing and extracting submatches from regular expressionsExplains transformations from regex to optimized automaton with practical optimizations

Latest Papers

What's happening recently
View more

This work addresses the problem of explaining why a finite automaton makes a specific decision on a given input word and how its output can be altered through minimal modifications. We formally define the notion of a minimal feature explanation for automaton decisions as the smallest set of critical input symbols responsible for the current acceptance or rejection outcome. To compute all such minimal explanations exactly, we propose an efficient algorithm that integrates formal methods, automata theory, and combinatorial optimization. Experimental evaluation demonstrates that our approach scales well in complex scenarios and consistently produces unbiased, accurate minimal explanations, thereby providing rigorous interpretability guarantees for automaton-based decisions.

decision reasoningexplanationfinite automata

This study addresses the challenges of distributional distortion and computational inefficiency in text generation constrained by nondeterministic finite automata (NFAs) by proposing the NFA-LM engine. Grounded in a hidden Markov model (HMM) formulation, this approach reduces the generation process to polynomial time complexity. Furthermore, by incorporating a fully polynomial-time randomized approximation scheme (FPRAS) based on #NFA counting, it establishes the first efficient constrained generation framework equipped with rigorous theoretical error bounds. Experimental results demonstrate that the proposed engine efficiently generates high-quality constrained text while providing reliable theoretical guarantees on approximation error, thereby effectively reconciling computational efficiency with theoretical rigor.

#NFAconstrained generationhard constraints

This study addresses the challenge of efficiently generating personalized corrective feedback from student-submitted finite automata under pedagogical constraints. To this end, this work proposes a novel tree-encoding technique that, for the first time, finitely represents the complete set of automaton corrections recognizing a given regular language as a regular tree language. Furthermore, a filtering mechanism is introduced to select correction candidates satisfying specific instructional requirements. By overcoming the computational bottlenecks inherent in traditional correction enumeration, the proposed approach enables the automated derivation and extraction of high-quality, personalized pedagogical feedback from erroneous student submissions. Consequently, this research establishes both a theoretical foundation and a practical toolset for advancing intelligent tutoring systems in formal language education.

Automata CorrectionDeterministic Finite AutomataDidactic Feedback

Current large language models (LLMs) ensure syntactic and constraint validity in structured generation but suffer from severely limited output diversity. To address this, we propose an automaton-guided generation mechanism that leverages historical state-transition trajectories—extracted during structured decoding—to dynamically steer the model toward under-explored structural patterns. By tightly integrating automata theory with LLM decoding, our method enhances structural and semantic diversity without compromising validity or inference efficiency. Experimental evaluation on open-source library test-case generation demonstrates a 27.4% improvement in diversity metrics—including structural coverage and semantic dissimilarity—while maintaining a 98.6% compliance rate with syntax and domain constraints. This confirms the method’s effectiveness and practical applicability for diverse, valid structured generation.

Enhancing diversity in automaton-based structured generation for LLMsOvercoming limited output diversity in structured generation methodsSteering LLMs toward novel structural patterns using automata history

Inference of Deterministic Finite Automata via Q-Learning

Oct 20, 2025
EH
Elaheh Hosseinkhani
🏛️ Universität zu Lübeck

This paper addresses the passive inference of deterministic finite automata (DFA), introducing Q-learning to this task for the first time and establishing a novel reinforcement learning (RL)-based paradigm for formal language learning. Methodologically, it reinterprets the Q-function semantics as a state-transition function over a finite domain, thereby enabling a rigorous mapping between sub-symbolic learning and symbolic systems; it further designs a tailored reward scheme, models discrete state-action spaces, and incorporates convergence guarantees to automatically induce DFA structure from input/output sequences. Experiments on multiple benchmark datasets successfully recover exact DFAs, demonstrating both effectiveness and robustness. The work extends the applicability of Q-learning beyond traditional RL domains and constructs an interpretable semantic bridge between reinforcement learning and formal language theory—offering a principled pathway toward the integration of symbolic and sub-symbolic AI.

Adapting reinforcement learning for automaton transition functionsBridging sub-symbolic learning with symbolic representationsUsing Q-learning for passive DFA inference

Hot Scholars

GD

Giuseppe De Giacomo

University of Oxford & Sapienza Università di Roma
Artificial IntelligenceData ManagementKnowledge RepresentationAutomated Planning
DH

Daniel Hausmann

University of Liverpool
Fixpoint theoryModal logicGame theory
NP

Nir Piterman

Professor in Computer Science, University of Gothenburg and Chalmers University of Technolog, Sweden
VerificationAutomataLogicGames
MC

Maxime Cordy

University of Luxembourg
Artificial IntelligenceMachine Learning SecurityTesting and VerificationSoftware Engineering
KV

Keyon Vafa

Harvard University
Machine learning