Automating Formal Verification with Reinforcement Learning and Recursive Inference

📅 2026-05-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses key limitations of large language models in automated formal verification—namely, scarce training data, difficulty adhering to strictly checkable specifications, and a tendency to exploit weak specifications through heuristic shortcuts. To mitigate these issues, the authors propose a recursive reasoning framework that integrates Group Relative Policy Optimization (GRPO) with verifier-guided feedback. The approach features multi-round reinforcement learning, verifier-driven reasoning scaffolds, and mechanisms for subgoal decomposition and proof revision. This methodology substantially reduces specification misuse and significantly improves correctness rates: on a refined Dafny benchmark, verification success rises from 9.7% to 31.1%; in Lean, combining proof revision boosts VeriCoding pass rates from 46.2% to 69.2%, and seven previously unsolved tasks in VERINA are successfully resolved.
📝 Abstract
Automated formal verification remains challenging for large language models because data for proof assistants and verification-aware languages is scarce, and correctness depends on satisfying precise machine-checkable specifications rather than producing plausible code. This thesis studies how verifier environments can improve LLM generation of verified programs and proofs through reinforcement learning from verifiable rewards (RLVR) and verifier-guided inference-time search. First, we train open-source models in Dafny with RLVR using Group Relative Policy Optimization (GRPO) and related variants, assembling generated candidates into complete programs and scoring them with compiler and verifier outcomes. Initial experiments on an APPS-derived Dafny dataset increased verified reward from 2.2% to 58.1%, but revealed specification hacking, where models exploit weak formal specifications instead of implementing the intended solutions. After filtering underspecified and vulnerable tasks, multi-turn RLVR on the refined benchmark improves the verified pass rate from 9.7% to 31.1%. Second, we develop a verifier-guided inference scaffold in Lean that treats proof generation as structured search over decomposed subgoals, verifier feedback, diagnostics, and repair. With a fixed base model, the full scaffold with proof reviser improves pass rate on an initial VeriCoding pilot set from 46.2% under direct repair to 69.2%. On the larger VERINA dataset, whole-task decomposition plus proof reviser solves 7 of 42 previously unsolved tasks. We also introduce Dalek-Bench, a repository-scale Lean benchmark derived from the Rust $\texttt{curve25519-dalek}$ verification project; preliminary results remain weak, indicating that stronger progress evaluation and task-specific tool-use policies are still needed.
Problem

Research questions and friction points this paper is trying to address.

formal verification
large language models
specification hacking
verifier-guided inference
reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning from Verifiable Rewards (RLVR)
Verifier-Guided Inference
Formal Verification
Proof Search
Large Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Max Tan
Department of Electrical Engineering and Computer Science