GrammarRL: Effective Grammar-Constrained Decoding via Reinforcement Learning

๐Ÿ“… 2026-09-30
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the semantic quality degradation and prohibitive computational costs of beam search in grammar-constrained generation by proposing a reinforcement learning framework that requires no annotated data. Built upon Llama large language models, the method introduces a bidirectional self-supervised reward mechanism combined with the RLOO optimization objective to jointly enhance syntactic compliance and semantic fidelity. Experimental results demonstrate that the proposed approach yields an average improvement of 9.8 BLEU points, with gains reaching up to 22.8 BLEU. It matches or surpasses beam search in performance while substantially reducing inference overhead, effectively balancing grammatical validity with generation quality.
๐Ÿ“ Abstract
Grammar-constrained generation guarantees syntactic validity, but can substantially degrade semantic quality when the model's preferred outputs are poorly aligned with the imposed grammar. This trade-off is particularly severe when the prompt is underspecified or the model has limited instruction-following ability. Beam search can partially mitigate these failures by exploring multiple valid sequences, but its computational cost grows with beam width, while sequence-level probability is only an imperfect proxy for semantic quality. We introduce GrammarRL, a label-free reinforcement learning method that adapts language models to grammar constraints without requiring annotated data. GrammarRL optimizes the model using two complementary self-supervised rewards derived from its own likelihoods: a direct reward, measuring how likely the constrained output is given the input, and a reverse reward, measuring how well the input can be reconstructed from the generated output. We optimize these rewards with a Reinforce Leave-One-Out (RLOO) objective over groups of grammar-constrained rollouts, augmented with the top-1 beam-search hypothesis and regularized towards a frozen base model. We evaluate GrammarRL on sign language gloss translation, hierarchical text classification, and named entity recognition using Llama models ranging from 1B to 8B parameters. GrammarRL consistently outperforms constrained greedy decoding, with an average improvement of 9.8 points and gains of up to 22.8 BLEU. It matches or outperforms beam search on two of the three tasks while preserving greedy-decoding inference cost. Ablations further show that the two rewards are complementary: either reward alone can underperform the untrained baseline, whereas their combination consistently improves upon it.
Problem

Research questions and friction points this paper is trying to address.

Grammar-constrained decoding
Semantic quality degradation
Constrained text generation
Reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Grammar-Constrained Decoding
Reinforcement Learning
Self-Supervised Rewards
RLOO
Label-Free
๐Ÿ”Ž Similar Papers
No similar papers found.