Score
Designs and builds automated systems that synthesize executable code from specifications and integrate formal verification (theorem provers, proof validation) into the generation loop to ensure the produced implementations satisfy formal properties. These systems additionally perform automated self-debugging and repair or rejection of counterexample-bearing candidates, apply code optimization, and analyze codebases and code-similarity metrics to guide selection, adaptation, and fallback implementations.
This work addresses the limited adoption of formal verification, which often requires expert-written annotations such as preconditions, postconditions, and loop invariants. To overcome this barrier, the authors propose a novel approach that leverages large language models (LLMs) in conjunction with assertions from test cases as static oracles to automatically generate Dafny verification annotations from code annotated with natural language comments. The method features an iterative refinement process guided by verifier feedback over multiple rounds and uniquely integrates multi-model LLM collaboration with a closed-loop verifier feedback mechanism. A VS Code plugin was developed to support practical deployment. Evaluated on 110 Dafny programs, the approach achieves a 98.2% annotation correctness rate within at most eight repair iterations. Empirical results highlight that proof-assistant-style annotation remains a key challenge for LLMs, while user feedback on the plugin was notably positive.
Existing large code models struggle to generate executable intermediate formal specifications, limiting precise verification and repair of program behavioral errors. This work proposes SpecCoder, a novel framework that focuses on generating executable inline assertions at critical program locations, thereby transforming static annotations into verifiable evidence. SpecCoder employs verification-guided training, fine-tuning the Qwen2.5-Coder series models using correct programs, behavioral mutants, and multi-round specification refinement trajectories. Evaluated on the HumanExec benchmark, SpecCoder substantially improves the correctness (+55.8%), completeness (+358.1%), and assertion validity (+26.6%) of inline specifications, significantly enhancing program verification and repair capabilities.
This work proposes a novel paradigm that bridges the long-standing divide between testing and formal verification in traditional software validation, enabling them to synergistically enhance both efficiency and quality. Grounded in Design by Contract, the approach leverages the counterexample generation capability of SMT solvers to transform formal verification tools into an integrated engine for automated testing and repair. Within a unified framework, the method simultaneously achieves three key objectives: automatic generation of test cases for faulty programs, construction of regression test suites with full coverage for correct programs, and correctness-guaranteed program repair. This represents the first integration of verification, testing, and repair into a single cohesive methodology.
This work proposes a human-AI collaborative paradigm for formal software specification that mitigates the traditional barriers to industrial adoption—namely, the notational complexity and high expertise threshold—while preserving the benefits of early error detection and explicit invariants. The approach employs an intermediate language blending natural language with lightweight LaTeX mathematical notation, enabling AI-assisted review, refinement, and code generation. Crucially, it distinguishes between components requiring rigorous formalization and those amenable to flexible treatment. By deeply integrating AI into the specification authoring and verification workflow, this method achieves “correct-by-construction” development in a case study on organizational knowledge growth simulation, significantly reducing costs while ensuring early validation and design correctness.
Automatically verifying C programs generated by large language models (LLMs) remains challenging due to their syntactic and semantic irregularities, which hinder formal verification. Method: This paper proposes SynVer—a novel framework that tightly integrates LLM-based program synthesis with formal verification. SynVer introduces verifiability-aware biasing mechanisms operating at both syntactic and semantic levels to guide LLMs toward generating verification-friendly code. It further incorporates separation logic (SL) specifications and the Verified Software Toolchain (VST) to enable end-to-end, fully automated verification—from specification to C implementation to machine-checked safety proofs. Results: Evaluated on diverse benchmarks covering basic coding tasks, SL assertions, and API specifications, SynVer significantly improves the automatic verification success rate of LLM-generated C programs. Empirical results demonstrate its scalability, robustness, and effectiveness in bridging the gap between neural code generation and rigorous formal assurance.
This work proposes the first agent-driven automated testing and repair framework targeting four classes of potential flaws—verified code, checker code, unverified code, and specifications—in verified compilers generated by coding agents. By leveraging the internal structure of compilers, the approach designs structure-aware, customized defect detection strategies and employs coding agents to automatically repair identified issues, with correctness ensured through formal verification and benchmark testing. Experimental evaluation on the Axon compiler demonstrates that the system effectively identifies and fixes defects without exhibiting reward-hacking behavior, thereby validating both the efficacy and safety of the proposed methodology.
This work addresses the high cost of manually writing formal specifications and the limitations of existing large language model (LLM)-based approaches that require white-box access to source code, thereby posing intellectual property and deployment constraints. The authors propose a black-box-driven method that leverages only test code and dynamic execution traces to generate candidate Java Modeling Language (JML) specifications via an LLM. These candidates are locally validated using bounded model checking, and an iterative feedback loop refines them based on verification outcomes. This approach is the first to enable fully automated formal specification generation without any access to the program’s internal structure. Evaluated on the SpecGenBench benchmark, it demonstrates that test-derived information effectively guides specification synthesis, while also highlighting critical challenges in checker compatibility and diagnostic feedback, substantially enhancing industrial applicability.
Large language models often introduce subtle, hard-to-detect bugs when generating complex software, compromising reliability. This work proposes the first fully automated, project-level code generation and verification framework based on an interactive theorem prover (ITP). The approach separates code with side effects into C++ while formalizing pure logical components in the ITP Rocq, where they are automatically verified and extracted for integration. When proofs fail, the concrete counterexample states guide an LLM agent to autonomously repair the code. In experiments, the system generated 1,859 lines of verified Rocq code and extracted 2,848 lines of C++ within 30 minutes, passing 265 unit tests and 12 hours of AFL++ fuzzing with zero crashes or hangs—outperforming Dafny’s backend, which failed to complete verification under identical conditions.
Automatically generating ACSL formal specifications for C programs is often hindered by insufficient semantic precision, heavy reliance on expert knowledge, and verification challenges. This work proposes a novel approach that integrates static analysis via Code Property Graphs (CPGs) with large language models (LLMs). By leveraging CPGs to extract key semantic features—such as arithmetic operations and loop structures—the method constructs structured prompts that deeply embed static analysis into LLM prompt engineering, enabling the generation of verifiable specifications enriched with runtime error prevention constraints. A closed-loop feedback mechanism with the Frama-C/WP verifier iteratively refines specification quality. Experiments on 604 C programs demonstrate a 98% specification generation success rate and a 96% complete proof rate, representing a 24.7%–51.7% improvement in complete proof rates over a pure code-prompting baseline across four mainstream LLMs.