Codoku: Renewable Program-Reasoning Challenges for Frontier Coding Agents

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of existing benchmarks to code-execution circumvention and data contamination by constructing a renewable program reasoning benchmark. This benchmark requires models to complete program fragments satisfying both static and dynamic constraints, thereby eliminating tool dependency. Methodologically, it introduces semantic concretization generation to construct novel puzzles with witnesses on demand, ensuring solvability. The solving process integrates control flow graph constraints, execution path verification, and an exponential sparse-space search algorithm. Experimental evaluations across five frontier large language models yield an average accuracy of approximately 50%, demonstrating that the proposed benchmark effectively assesses genuine program reasoning capabilities.
📝 Abstract
Existing program-reasoning benchmarks ask large language models to predict a program's behavior on a given input. Coding agents break two assumptions on which these benchmarks rest: an agent can recover the answer by executing the program instead of reasoning about it, and fixed task sets drawn from existing programs are increasingly exposed to contamination, yet costly to renew. We introduce Codoku (code sudoku), a renewable benchmark in which a solver fills typed cells in a partial program to satisfy global static and dynamic constraints, such as a prescribed control-flow graph and execution path. Because a partial program cannot be executed and valid fillings are sparse in an exponentially large space of interdependent choices, neither tool use nor enumeration can substitute for program reasoning. Puzzles are synthesized from scratch via semantic reification, so fresh puzzles of controllable complexity can be generated on demand, each with a witness that guarantees solvability. We evaluate five frontier models on 300 puzzles through a coding agent free to use any tool within a fixed budget. Small puzzles already challenge open-weight models, whereas even proprietary models solve only about half of the large ones. Codoku thus offers a renewable testbed for program reasoning that can keep pace with rapidly improving coding agents. GitHub: https://github.com/connglli/Codoku.
Problem

Research questions and friction points this paper is trying to address.

program reasoning
coding agents
benchmark contamination
renewable benchmark
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Program Reasoning
Renewable Benchmark
Semantic Reification
Coding Agents
Partial Program Synthesis
🔎 Similar Papers
2024-02-08International Conference on Machine LearningCitations: 6