🤖 AI Summary
This study addresses the challenges of distributional distortion and computational inefficiency in text generation constrained by nondeterministic finite automata (NFAs) by proposing the NFA-LM engine. Grounded in a hidden Markov model (HMM) formulation, this approach reduces the generation process to polynomial time complexity. Furthermore, by incorporating a fully polynomial-time randomized approximation scheme (FPRAS) based on #NFA counting, it establishes the first efficient constrained generation framework equipped with rigorous theoretical error bounds. Experimental results demonstrate that the proposed engine efficiently generates high-quality constrained text while providing reliable theoretical guarantees on approximation error, thereby effectively reconciling computational efficiency with theoretical rigor.
📝 Abstract
Constrained generation aims to sample from language models (LMs) conditioned on hard constraints. Existing constrained-generation techniques for nondeterministic finite automaton (NFA) constraints either distort the distribution or sacrifice efficiency. Theoretically, this task reduces to counting the length-$n$ sequences accepted by an NFA (#NFA), and the exact #NFA problem is #P-complete. Recent work has shown that #NFA admits a fully polynomial randomized approximation scheme (FPRAS). Inspired by this result, we propose NFA-LM, a polynomial-time engine for NFA-constrained generation with theoretical guarantees under mild assumptions. Experiments show that NFA-LM efficiently generates high-quality outputs with theoretically bounded approximation error.