🤖 AI Summary
This study addresses the limitations of existing benchmarks, which lack evaluation of end-to-end self-generated specification–code verification linkages and fail to cover multilingual real-world scenarios. To bridge this gap, we propose VeriCodeBench, a comprehensive benchmark, alongside CodeNova, a novel framework that pioneers an end-to-end self-specifying verifiable code generation paradigm. In this approach, large language models autonomously generate constraint-guided formal specifications and corresponding code, while verifier feedback drives targeted repair, establishing a full closed-loop pipeline across C, Java, Rust, and Python. Experimental results demonstrate that our method yields significant improvements across all evaluated metrics. Notably, Claude Sonnet 5 achieves state-of-the-art performance under the self-specification protocol. Furthermore, our analysis reveals that specification generation remains the primary bottleneck in current verifiable code synthesis pipelines.
📝 Abstract
Large language models (LLMs) may generate unreliable code on corner cases missed by testing, while formal verification can provide machine-checkable guarantees. Recently, researchers have proposed several benchmarks to evaluate the capabilities of LLMs in generating formally verifiable code, where LLMs need to formulate formal specifications, generate the corresponding code, and verify its correctness. However, existing benchmarks have two key limitations: (I) They primarily evaluate specification and code generation stage-wise, with code generation typically conditioned on an oracle specification. This setup overlooks whether strong stage-wise performance translates into end-to-end success. (II)They mainly focus on a single proof-oriented language and mathematically structured tasks, offering limited coverage of tasks common in software development. In this paper, we introduce VeriCodeBench, a benchmark for self-spec verifiable code generation, where the LLM relies solely on its own generated specification and code throughout the entire process. VeriCodeBench contains 400 language-native problems across C, Java, Rust, and Python, covering practical concerns in software development. We evaluate specification coverage, code validity, and joint problem-level success. We further introduce CodeNova to enhance the capabilities of LLMs in self-spec verifiable code generation. CodeNova makes requirements explicit through constraint-guided specification and uses verifier feedback to guide targeted implementation repairs. Experimental results reveal that self-generated specifications remain a major bottleneck, while providing more sophisticated specifications may not necessarily lead to higher verification success rates. CodeNova substantially improves performance across all evaluation metrics, enabling Claude Sonnet 5 to achieve the strongest results under the self-spec protocol.