GRID: Grammar-Railed Decoding for Enterprise SQL Generation

📅 2026-07-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Enterprise-grade SQL generation demands strict adherence to syntactic correctness, role- and schema-compliance, provable guarantees, low latency, and full auditability—requirements that general-purpose large language models (LLMs) struggle to satisfy. This work proposes GRID, a grammar-constrained decoding engine that steers LLM outputs into the valid prefix space of an LALR(1) grammar, embedding role-based access control directly within the grammar specification. GRID dynamically generates decoding masks using parser states—comprising lexical scanner states and the LALR(1) parsing stack—thereby providing formal guarantees of soundness, completeness, termination, and near-constant per-token overhead. Experiments show that GRID improves execution accuracy by 13 percentage points on Spider; with a single repair pass, a 7B-parameter model achieves 94.5% executable rate. Masking incurs a median latency of only 3.6–6.7 microseconds and enables bit-level audit replay with 100% tamper detection.
📝 Abstract
Large language models can write SQL, but enterprise deployment demands more than plausible text: outputs must be syntactically valid, must respect per-role and per-schema policy, must carry provable (not best-effort) guarantees, must not slow down as generations grow, and must leave a compliance-grade record of every decision. We present GRID (Grammar-Railed Decoding), a grammar-constrained decoding engine that keys exact next-token masks on parser configurations (lexer scan state x LALR(1) stack) rather than on token sequences, and uses the incrementally advanced LALR(1) parser itself as a viable-prefix oracle. LLM tokens are bridged to grammar terminals by a byte-level trie walk with a context-independent/context-dependent split that makes cache-key soundness hold by construction. Role-based access control is compiled into the language: role projections subset the grammar's productions and schema lexicons restrict identifier terminals, so forbidden verbs and identifiers are unreachable at mask level. Four guarantees (soundness, completeness, termination, and near-constant per-token cost) are stated with explicit preconditions and each paired with a test or benchmark. Rust kernels bring the per-token mask to a 3.6-6.7 us median, ahead of llguidance at p50 and p90 on two tokenizers with zero false rejects; per-token guard cost is position-flat at n=16,000. On Spider, constrained decoding is worth +13 execution-accuracy points at 0.5B, and one checker-guided repair pass over the provably mask-unenforceable residue (column-level policy) lifts a 7B model to 94.5% executable. A hash-chained per-token audit trail replays bit-identically with 100% tamper detection. We state plainly what the mask cannot do (distribution faithfulness, column-level RBAC, non-LALR(1) languages) and where measured cost remains.
Problem

Research questions and friction points this paper is trying to address.

enterprise SQL generation
grammar-constrained decoding
role-based access control
compliance auditing
syntactic validity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Grammar-Constrained Decoding
LALR(1) Parsing
Role-Based Access Control
Incremental Parsing Oracle
Compliance-Auditable Generation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mohsen Arjmandi
evolutionID GmbH