Score
Designs, implements, and evaluates generation and decoding pipelines that constrain and guide natural-language outputs from language models so they conform to explicit restrictions (e.g., controlled natural-language subsets, restricted vocabularies, ontology- or schema-based formats, or machine-verifiable rule syntax). Builds and applies constraint and refinement techniques—such as guided decoding, n-gram or pattern constraints, linguistic refinements, and post-editing—to reduce inconsistency, repetition, and awkward phrasing and to improve fluency and coherence.
Existing constrained decoding approaches treat schemas solely as structural constraints, overlooking the potential influence of their linguistic formulation on large language model behavior. This work reframes structured generation as a multi-channel instruction problem, demonstrating that subtle adjustments to schema key wording can implicitly convey instructions to guide model outputs—without altering prompts or model parameters. We systematically reveal, for the first time, that schema phrasing serves as an effective implicit instruction channel. Furthermore, we find significant differences across model families in their sensitivity to prompt-level versus schema-level instructions, with their interaction exhibiting non-additive effects. Experiments show that Qwen substantially benefits from schema-level instructions in mathematical reasoning tasks, whereas LLaMA relies more heavily on prompt-level guidance, and combining both channels does not necessarily yield cumulative performance gains.
Large language models often generate semantically incorrect code in low-resource programming languages due to undefined variable references, invalid fields, or unsupported options. This work proposes a runtime environment-aware decoding mechanism that dynamically instantiates grammar fragments from an environment Γ, employs a region-based policy to select valid syntactic structures, and resolves open references by filling Γ-typed slots, thereby guaranteeing both syntactic well-formedness and semantic validity while enabling immediate feedback for newly declared constructs. We formalize environment-indexed grammars and their refinement order, prove their preservation of semantic correctness, and characterize the boundary of mask-enforceable properties. Experiments on TileLang, SQL, and P4 demonstrate that the gproj system eliminates phantom references with minimal overhead and substantially improves the semantic correctness of generated code.
This work addresses the low compliance and lack of systematic evaluation of constrained decoding techniques under realistic, complex constraints—particularly JSON Schema. We introduce JSONSchemaBench, the first large-scale structured generation benchmark comprising 10K real-world JSON schemas, and propose a multidimensional evaluation framework assessing compliance, constraint coverage, and output quality across six state-of-the-art methods. Our analysis reveals, for the first time, that compliance rates drop by over 40% for existing frameworks when handling nested, recursive, and conditional constraints. We further identify XGrammar and Outlines as achieving the best trade-off between inference efficiency and generation quality. The benchmark and evaluation suite are fully open-sourced, filling a critical gap in systematic evaluation for structured generation and advancing constrained decoding toward higher reliability and stronger generalization.
Large language models (LLMs) struggle to reliably adhere to *soft constraints*—semantically rich, human-intended requirements that lack automated verifiability—hindering their trustworthy deployment in real-world applications. To address this, we propose the first systematic framework for soft constraint adherence: (1) a constraint-aware curriculum learning paradigm that incrementally increases semantic complexity during training; (2) a fully automated, annotation-free pipeline for generating high-quality soft constraint data; and (3) a constraint-driven prompt optimization and evaluation protocol. Our approach achieves significant improvements in constraint adherence across multiple soft-constraint benchmarks. Ablation studies confirm the critical roles of both the curriculum strategy and data quality. All generated data, implementation code, and evaluation protocols are publicly released to foster reproducibility and further research.
Existing approaches to natural language generation under strict hard constraints—such as the RADNER rules—exhibit significant limitations in both constraint expressivity and adherence. Method: This paper introduces a “constraint-first” framework that systematically integrates constraint programming (CP) into NLP text generation for the first time. It formalizes generation as a discrete combinatorial optimization problem, jointly encoding linguistic features (e.g., n-grams, syllables, character counts) and declarative constraints. A large language model (LLM) then ranks candidate outputs via perplexity-based scoring to identify the optimal solution. Crucially, the method requires no fine-tuning or prompt engineering. Contribution/Results: The framework achieves fully automatic compliance with extremely stringent syntactic and structural constraints. Empirical evaluation in clinical and vision-science domains demonstrates robust generation of large-scale, constraint-satisfying sentences—even under “unreasonably strong” constraints—establishing a novel paradigm for hard-constraint text generation.
Constrained decoding in code generation often suffers from misalignment between the constraint enforcer (e.g., a type system), large language models, and the target programming language (e.g., TypeScript), leading to degraded functional correctness. This work is the first to demonstrate that when constraint enforcers are incomplete or unsound, constrained decoding inadvertently steers models toward low-probability program regions, significantly reducing correctness—sometimes performing worse than unconstrained decoding, which achieves up to a 97% lower error rate in certain scenarios and incurs fewer timeouts. The study systematically evaluates decoding behaviors across seven large language models, two programming languages, and two classes of constraint enforcers on three benchmarks, and provides quantitative design principles for developing effective constraint mechanisms.
This work addresses the challenge of generating executable Cypher queries from natural language, a task often hindered by outputs that violate syntactic validity or database schema consistency. The authors propose a training-free, test-time structural constraint filtering framework that applies, during inference, a multi-stage post-processing pipeline comprising confidence scoring, context-free grammar validation, and graph database schema consistency checking. This approach explicitly disentangles and quantifies the distinct contributions of syntactic and schema-level constraints to query quality. Experimental results demonstrate significant improvements in both syntactic correctness and execution accuracy across two instruction-tuned models. Specifically, grammar-based filtering markedly enhances syntactic compliance, while schema-aware filtering further boosts semantic correctness, albeit at the cost of reduced coverage under stringent constraints.
The continuous evolution of large language models induces prompt behavior drift, undermining the stability and controllability of traditional prompt engineering. To address this, this work proposes the Natural Language Declarative Prompting (NLD-P) framework, which reconceptualizes prompt design as a declarative governance approach. NLD-P modularly abstracts source specifications, constraint logic, task content, and post-generation evaluation, encoding control structures entirely in natural language without external code orchestration. This framework elevates prompt engineering to a system-level governance paradigm, enabling non-technical users to achieve interpretable and stable prompt management amid model evolution. The study defines minimal compliance criteria for NLD-P, validates its applicability across model versions, and introduces a human-in-the-loop verification mechanism alongside model-dependent pattern acceptability analysis, offering a novel pathway for prompt control in dynamic model environments.
This work addresses the challenge of efficiently avoiding multiple hard constraints—such as sensitive terms and personally identifiable information (PII)—as well as regular expression patterns during large language model (LLM) generation. Conventional automaton-based approaches often suffer from state-space explosion and high computational overhead. To overcome these limitations, the authors propose NCO decoding, a lightweight online pattern-matching strategy that dynamically enforces finite hard and regular-expression constraints without constructing large automata. NCO is the first method to enable online decoding control under multiple hard and regex constraints, seamlessly integrating with beam search and various sampling strategies. It further incorporates a soft masking mechanism to probabilistically suppress prohibited content. Experiments demonstrate that NCO significantly reduces violation rates in PII and profanity suppression tasks while preserving generation quality and inference efficiency.
This work addresses the frequent lack of semantic validity in code generated by large language models (LLMs) for software engineering tasks. To this end, it introduces a projection decoding framework that, for the first time, treats graph-based representations as first-class citizens alongside textual sequences during generation. The framework incrementally constructs partial graph structures in parallel with token prediction, directly embedding domain-specific semantics into the decoding process. This integration enables uncertainty modeling, incremental semantic validation, and provable correctness guarantees, thereby establishing a verifiable foundation for LLM-driven software engineering automation. Experimental results demonstrate that the proposed approach significantly improves the semantic validity of generated artifacts in program synthesis tasks.