🤖 AI Summary
Large language models (LLMs) frequently generate syntactically valid but type-incorrect code, leading to compilation failures; existing constrained decoding methods address only syntactic constraints and lack semantic type awareness. This paper introduces Type-Constrained Decoding, the first approach to deeply integrate a formal type system into LLM decoding. It constructs a type-aware prefix automaton grounded in type inference and inhabitation search, enabling sound and efficient decoding under type constraints. The method is rigorously formalized for simply typed languages and successfully extended to TypeScript’s richer semantics. Experiments on HumanEval show over 50% reduction in compilation errors and significant gains in functional correctness. The technique demonstrates consistent improvements across diverse model scales—including state-of-the-art open-source models with >30B parameters—and delivers robust performance gains in code synthesis, translation, and repair tasks.
📝 Abstract
Large language models (LLMs) have achieved notable success in code generation. However, they still frequently produce uncompilable output because their next-token inference procedure does not model formal aspects of code. Although constrained decoding is a promising approach to alleviate this issue, it has only been applied to handle either domain-specific languages or syntactic language features. This leaves typing errors, which are beyond the domain of syntax and generally hard to adequately constrain. To address this challenge, we introduce a type-constrained decoding approach that leverages type systems to guide code generation. We develop novel prefix automata for this purpose and introduce a sound approach to enforce well-typedness based on type inference and a search over inhabitable types. We formalize our approach on a simply-typed language and extend it to TypeScript to demonstrate practicality. Our evaluation on HumanEval shows that our approach reduces compilation errors by more than half and increases functional correctness in code synthesis, translation, and repair tasks across LLMs of various sizes and model families, including SOTA open-weight models with more than 30B parameters.