๐ค AI Summary
This study addresses the high memory consumption and low decoding efficiency caused by automaton constraints during structured reasoning in language models. To overcome these limitations, this work proposes an efficient inference engine grounded in path-sparse graph theory. Methodologically, it surpasses the space bounds of Savitchโs theorem by designing a reachability decision algorithm with O(logยฒn/loglogn) complexity, enabling table-free, real-time GPU mask recomputation and parallel recursive decoding. Experimental results demonstrate that the proposed approach reduces grammar memory requirements from gigabytes to megabytes. Furthermore, it achieves 1.2โ1.3ร faster extraction on Apple M2 Pro hardware, accelerates tool-calling agents by 2.5ร, and supports concurrent multi-grammar execution.
๐ Abstract
Language models are more and more often asked for structured output: JSON that follows a schema, or a tool call with typed arguments. A small machine, an automaton, enforces the format by forbidding the tokens that would break it. We observe that this machine has a rare property: from any of its states, each token leads along exactly one path. Graphs in which only a few paths join any two points are a classical object of complexity theory, and our theoretical result settles an open question about them: one can decide whether such a graph connects two points while verifying that it really has few paths, with very little memory. Precisely, the problem lies in the classes ReachUL, LOGDCFL, C=L and SC2, and needs only O(log2 n/ log log n) space, below the classical O(log2 n) of Savitch's theorem. The constructions behind the proofs become an inference engine: text the format forces is written without running the model, the mask is recomputed on the GPU without any table, recursive formats use a small stack, every output stays valid under a token limit, and independent fields are decoded in parallel and verified. On one 16 GB Apple M2 Pro with Qwen3.5-2B and 4B, against MLX with llguidance, the standard setup for this hardware, schema-constrained extraction finishes 1.2- 1.3x sooner with the same answers, a grammar costs 3 MB instead of up to 1.5 GB, one server holds sixteen grammars where tables run out of memory, and sixteen tool-calling agents finish 2.5x sooner.