constrained decoding

Designs and implements decoding algorithms and runtime systems that enforce lexical, regex, graph, or syntax constraints during sequence generation from probabilistic sequence models, so generated token sequences provably or heuristically satisfy specified constraints. Work includes building trie-based and dynamic-prefix decoders, Viterbi-style constrained search, token-filtering/sampling hooks for forbidden-word bans, and engineering trade-offs between constraint compliance, efficiency, and output quality.

constraineddecoding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.12
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Type-Constrained Code Generation with Language Models

Apr 12, 2025
NM
Niels Mundler
🏛️ ETH Zurich | UC Berkeley

Large language models (LLMs) frequently generate syntactically valid but type-incorrect code, leading to compilation failures; existing constrained decoding methods address only syntactic constraints and lack semantic type awareness. This paper introduces Type-Constrained Decoding, the first approach to deeply integrate a formal type system into LLM decoding. It constructs a type-aware prefix automaton grounded in type inference and inhabitation search, enabling sound and efficient decoding under type constraints. The method is rigorously formalized for simply typed languages and successfully extended to TypeScript’s richer semantics. Experiments on HumanEval show over 50% reduction in compilation errors and significant gains in functional correctness. The technique demonstrates consistent improvements across diverse model scales—including state-of-the-art open-source models with >30B parameters—and delivers robust performance gains in code synthesis, translation, and repair tasks.

Addressing typing errors beyond syntactic constraints in code generationEnhancing functional correctness in code synthesis and translation tasksReducing uncompilable code output from large language models

This work addresses the high latency of large language models in context-free grammar (CFG)-constrained decoding, which stems from the need to traverse the full vocabulary at each generation step, rendering complex grammars impractical. To overcome this limitation, the authors propose CFGzip, a novel offline method that compresses the token search space by integrating CFG analysis with semantic token clustering to construct a compact yet complete subset of valid tokens. This approach preserves generation correctness while drastically reducing the search scope. CFGzip seamlessly integrates with existing grammar-constrained decoding engines and, in standard evaluations, achieves up to a 7.5× speedup and reduces latency by two orders of magnitude, substantially enhancing the practicality and scalability of generating text under complex CFG constraints.

constrained decodingcontext-free grammardecoding overhead

This work addresses the performance bottleneck—termed the “cardinality wall”—faced by existing grammar-constrained decoding methods for large language models when generating structured outputs conforming to predefined schemas, particularly when the candidate set size is large (K ≥ 300). The authors propose a novel automaton-based mechanism that integrates prefix trees (Tries) with the Aho-Corasick multi-pattern matching algorithm. By leveraging shared prefixes, bounded depth, and known cardinality of the candidate set, the method precomputes valid token masks at each trie node, enabling stateless and highly efficient constrained decoding. This approach, the first to combine Trie and Aho-Corasick techniques for decoding constraints, achieves sub-100ms compilation time even at K = 10,000, offering 2–6.5× faster compilation and 7× lower per-token latency (0.65 μs vs. 5.8 μs) than XGrammar. At batch size 256, it attains an end-to-end throughput of 219 requests/second—a 29× improvement—while guaranteeing 100% output validity.

cardinality wallconstrained decodinggrammar compilation

Large language models (LLMs) struggle to simultaneously ensure syntactic correctness and distributional fidelity when generating highly structured outputs (e.g., code, mathematical expressions, markup). While grammar-constrained decoding (GCD) guarantees syntactic compliance, it severely distorts the model’s original conditional distribution. This work formally defines the *grammar alignment* problem and proposes ASAP—a novel decoding framework that strictly enforces context-free grammar constraints while provably preserving the LLM’s conditional output distribution. ASAP integrates adaptive sampling, approximate expected future state estimation, grammar-guided over-approximation of prefix feasibility, and context-aware constraint propagation. Experiments on code generation and structured NLP tasks demonstrate that ASAP achieves significantly higher likelihood under the original LLM distribution compared to baselines, while maintaining 100% syntactic validity.

Constrained decoding distorts LLM probability distributionsLLMs struggle with structured output generationNeed algorithm ensuring grammar compliance and distribution alignment

Latest Papers

What's happening recently
View more

This work addresses the challenge that large language models often generate code violating the syntax of domain-specific languages (DSLs) when invoking external services, a problem exacerbated by the absence of context-free grammars for third-party DSLs needed for syntax-constrained decoding. To overcome this, the authors propose Autogrammar, an agent that uniquely integrates Kripke structures with language models, enabling declarative control of agent behavior via linear temporal logic and automatically inducing DSL grammars from documentation and execution feedback—eliminating the need for manual grammar engineering. Evaluated on three real-world DSLs, the learned grammars achieve near-perfect precision (≈100%) on unseen data, significantly outperforming baseline methods in end-to-end task accuracy, matching or exceeding human-crafted grammars in 80% of tasks, while accelerating inference by 3.8×.

context-free grammardomain-specific languagegrammar-constrained decoding

Large language models often generate semantically incorrect code in low-resource programming languages due to undefined variable references, invalid fields, or unsupported options. This work proposes a runtime environment-aware decoding mechanism that dynamically instantiates grammar fragments from an environment Γ, employs a region-based policy to select valid syntactic structures, and resolves open references by filling Γ-typed slots, thereby guaranteeing both syntactic well-formedness and semantic validity while enabling immediate feedback for newly declared constructs. We formalize environment-indexed grammars and their refinement order, prove their preservation of semantic correctness, and characterize the boundary of mask-enforceable properties. Experiments on TileLang, SQL, and P4 demonstrate that the gproj system eliminates phantom references with minimal overhead and substantially improves the semantic correctness of generated code.

constrained decodingdomain-specific languagesghost references

Enterprise-grade SQL generation demands strict adherence to syntactic correctness, role- and schema-compliance, provable guarantees, low latency, and full auditability—requirements that general-purpose large language models (LLMs) struggle to satisfy. This work proposes GRID, a grammar-constrained decoding engine that steers LLM outputs into the valid prefix space of an LALR(1) grammar, embedding role-based access control directly within the grammar specification. GRID dynamically generates decoding masks using parser states—comprising lexical scanner states and the LALR(1) parsing stack—thereby providing formal guarantees of soundness, completeness, termination, and near-constant per-token overhead. Experiments show that GRID improves execution accuracy by 13 percentage points on Spider; with a single repair pass, a 7B-parameter model achieves 94.5% executable rate. Masking incurs a median latency of only 3.6–6.7 microseconds and enables bit-level audit replay with 100% tamper detection.

compliance auditingenterprise SQL generationgrammar-constrained decoding

Constrained decoding in code generation often suffers from misalignment between the constraint enforcer (e.g., a type system), large language models, and the target programming language (e.g., TypeScript), leading to degraded functional correctness. This work is the first to demonstrate that when constraint enforcers are incomplete or unsound, constrained decoding inadvertently steers models toward low-probability program regions, significantly reducing correctness—sometimes performing worse than unconstrained decoding, which achieves up to a 97% lower error rate in certain scenarios and incurs fewer timeouts. The study systematically evaluates decoding behaviors across seven large language models, two programming languages, and two classes of constraint enforcers on three benchmarks, and provides quantitative design principles for developing effective constraint mechanisms.

alignment problemcode generationconstrained decoding

Existing grammar-constrained decoding methods suffer from high computational overhead under large vocabularies, severely limiting the throughput efficiency of large language models in generating structured text. This work proposes PSC (Parsing Stack Classifier), the first approach to achieve vocabulary-size-independent mask computation complexity. By preprocessing all token-level syntactic acceptance conditions into a unified parsing stack classifier, PSC enables the generation of a complete constraint mask at each decoding step through a single stack state check. Integrating context-free grammars, parsing stack modeling, and an efficient classifier design, the method achieves speedups of up to 700× and 30× on programming language and JSON generation tasks, respectively, with end-to-end throughput approaching that of unconstrained decoding.

computational overheadgrammar-constrained decodinglarge language models

Hot Scholars

AH

Ahmed Hareedy

Assistant Professor, EEE Department, Middle East Technical University
Coding TheoryInformation TheoryOptimizationData Storage
JL

Jia Li

Assistant Professor, College of AI, Tsinghua University
Programming Language ProcessingFoundation ModelAI Agent
ZW

Zhihao Wen

Singapore Management University
Graph Neural NetworkLarge Language ModelParameter Efficient Fine-TuningMeta-learning
AS

Amrit Singh Bedi

Assistant Professor in Department of Computer Science, University of Central Florida, FL, USA
Reinforcement LearningAI AlignmentAI generated Text DetectionConvex and Non-convex optimization
MA

Mohammad Albinhassan

PhD Student, Imperial College London
Large Language ModelsReinforcement LearningNeuro-Symbolic AI