Improving Large Language Models for Code through Runtime Program-State Reasoning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inability of large language models to explicitly reason about runtime program states in code-related tasks by incorporating state reasoning into the post-training pipeline. Building upon the Qwen3.5-9B foundation model, this work introduces two reasoning tasks—defective input-output analysis and precondition-postcondition verification—and employs a phased training strategy combining supervised fine-tuning with sequential reinforcement learning. This approach achieves, for the first time, a deep integration of runtime state reasoning into large model training. The resulting Comet-9B model substantially outperforms same-scale baselines on benchmarks such as SWE-bench Pro, attaining performance comparable to GPT-4o and GPT-5.2 agents. These findings demonstrate that explicit state reasoning can significantly enhance the software engineering capabilities of smaller-parameter models.
📝 Abstract
Large language models receive limited explicit training in reasoning about runtime program states. We study whether training models to reason about runtime program states improves downstream software-engineering capabilities. We introduce two complementary program-state reasoning tasks. Buggy input-output reasoning requires a model to generate a concrete input that exposes a behavioral difference between a buggy program and a hidden correct implementation and to predict the resulting execution behavior. Precondition-postcondition reasoning requires an agent to symbolically characterize a bug-triggering precondition, predict the expected postcondition, explain their causal connection, and instantiate this reasoning as an executable regression test. By incorporating these two tasks into a staged post-training pipeline, we develop Comet-9B, a 9B language model based on Qwen3.5-9B Base. We evaluate the resulting checkpoints on repository-level patch generation, regression-test generation, and security PoC generation. Adding both program-state reasoning tasks to supervised fine-tuning (SFT) on issue resolution improves success rates by 7.25 percentage points on SWE-bench Pro and 9.70 points on SWT-Bench Verified. Sequential reinforcement learning on the two tasks yields further gains of 7.25, 26.79, and 4.67 percentage points on SWE-bench Pro, SWT-Bench Verified, and CyberGym, respectively. Despite having only 9B parameters, Comet-9B achieves a score comparable to the reported GPT-5.2 result on SWE-bench Pro and matches the reported success rate of a GPT-4o-based agent on SWT-Bench Verified.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Runtime Program-State Reasoning
Software Engineering
Code Generation
Bug Detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Runtime Program-State Reasoning
Staged Post-Training Pipeline
Sequential Reinforcement Learning
Supervised Fine-Tuning
Large Language Models for Code
🔎 Similar Papers
2024-02-08International Conference on Machine LearningCitations: 6