Breaking the Space Barrier and its Application to Language Model Inference

๐Ÿ“… 2026-10-06
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the high memory consumption and low decoding efficiency caused by automaton constraints during structured reasoning in language models. To overcome these limitations, this work proposes an efficient inference engine grounded in path-sparse graph theory. Methodologically, it surpasses the space bounds of Savitchโ€™s theorem by designing a reachability decision algorithm with O(logยฒn/loglogn) complexity, enabling table-free, real-time GPU mask recomputation and parallel recursive decoding. Experimental results demonstrate that the proposed approach reduces grammar memory requirements from gigabytes to megabytes. Furthermore, it achieves 1.2โ€“1.3ร— faster extraction on Apple M2 Pro hardware, accelerates tool-calling agents by 2.5ร—, and supports concurrent multi-grammar execution.
๐Ÿ“ Abstract
Language models are more and more often asked for structured output: JSON that follows a schema, or a tool call with typed arguments. A small machine, an automaton, enforces the format by forbidding the tokens that would break it. We observe that this machine has a rare property: from any of its states, each token leads along exactly one path. Graphs in which only a few paths join any two points are a classical object of complexity theory, and our theoretical result settles an open question about them: one can decide whether such a graph connects two points while verifying that it really has few paths, with very little memory. Precisely, the problem lies in the classes ReachUL, LOGDCFL, C=L and SC2, and needs only O(log2 n/ log log n) space, below the classical O(log2 n) of Savitch's theorem. The constructions behind the proofs become an inference engine: text the format forces is written without running the model, the mask is recomputed on the GPU without any table, recursive formats use a small stack, every output stays valid under a token limit, and independent fields are decoded in parallel and verified. On one 16 GB Apple M2 Pro with Qwen3.5-2B and 4B, against MLX with llguidance, the standard setup for this hardware, schema-constrained extraction finishes 1.2- 1.3x sooner with the same answers, a grammar costs 3 MB instead of up to 1.5 GB, one server holds sixteen grammars where tables run out of memory, and sixteen tool-calling agents finish 2.5x sooner.
Problem

Research questions and friction points this paper is trying to address.

Language Model Inference
Structured Output
Space Complexity
Constrained Decoding
ReachUL
Innovation

Methods, ideas, or system contributions that make the work stand out.

constrained decoding
space complexity
structured output
GPU inference
automata theory
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.