Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过创建SWE-Flux基准来评估大语言模型在代码执行推理方面的能力,该基准包含480个基于真实Python仓库的执行实例。
📝 Abstract
Large language models (LLMs) are increasingly used in coding tasks, but their ability to reason about code execution remains unclear. Existing repository-level QA benchmarks mainly evaluate static code understanding and often rely on LLM-based evaluation, while execution-reasoning benchmarks are mostly limited to snippets or functions. We introduce SWE-Flux, a repository-level benchmark for dynamic execution reasoning containing 480 execution-grounded instances across 12 real Python repositories, with gold answers automatically harvested from instrumented test executions rather than written manually or judged by LLMs. The benchmark covers singletest and multi-test questions over control flow, loops, program state, dataflow, exceptions, and program invariants. Evaluating five LLMs shows that this task remains challenging. The best model achieves only 37% accuracy. Models perform better on localized behavior such as invariants, intra-procedural control flow, exceptions, and simple loops, but struggle with dataflow, inter-procedural execution, precise state reasoning, and suite-level aggregation. Finally, we show that the oracle-harvesting pipeline can generate fresh benchmark variants using input perturbation. It successfully harvests valid variants for almost 90% of the selected instances, and the resulting variants are substantially more challenging for the evaluated models.
Problem

Research questions and friction points this paper is trying to address.

Large language models
code execution reasoning
repository-level
dynamic behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

dynamic execution reasoning
repository-level benchmark
automatically harvested gold answers
input perturbation
🔎 Similar Papers
No similar papers found.