llm-assisted theorem proving

Designs and implements closed-loop systems that combine large language models and formal/verifier engines to generate, translate, and synthesize proofs: practitioners build pipelines in which an LLM proposes natural-language or formal proof steps, invariants, lemmas, or end-to-end proof scripts and a verifier or automated prover checks them and returns errors, counterexamples, or obligations. They then develop the iterative refinement logic that interprets verifier feedback, coordinates symbolic computation, explores proof strategies, and updates proposals until a machine-certified formalization or verified proof (and optionally a human-readable proof) is produced.

llm-assistedtheoremproving

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.27
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Current benchmarks for mathematical reasoning predominantly rely on answer matching, which fails to assess the logical correctness of solution processes. This work proposes a hybrid verification pipeline that integrates automated and interactive validation by leveraging structured prompting to guide large language models in generating verifiable solutions. The framework supports both formal and informal reasoning and interfaces with proof assistants such as Lean 4, enabling even small-scale models (≤8B parameters) to participate effectively in collaborative verification. Through a multi-agent architecture and advanced prompt engineering, the approach substantially reduces false positive rates. Experimental results demonstrate high verification accuracy across multiple datasets, and the codebase along with deployment guidelines has been publicly released.

BenchmarkingFalse PositivesLarge Language Models

A Case Study on the Effectiveness of LLMs in Verification with Proof Assistants

Aug 25, 2025
BB
Barış Bayazıt
🏛️ University of Toronto | Portland State University

This work investigates the effectiveness of large language models (LLMs) in assisting interactive theorem proving, particularly within real-world, industrial-scale formal verification tasks. Method: We conduct a systematic evaluation on two authentic verification projects—hs-to-coq and Verdi—using the Rocq proof assistant, employing both quantitative metrics (success rate, proof length, error rate) and qualitative analysis (tactic reasonableness, technical reusability). Contribution/Results: LLMs demonstrate strong capability in generating concise, high-quality formal proofs aligned with classical proof styles, scaling effectively from small to large proofs. Performance critically depends on external dependency information and contextual modeling quality, exhibiting marked heterogeneity across projects. While rare anomalous errors occur, they are substantially mitigated via context enhancement. To our knowledge, this is the first empirical study to characterize LLM capabilities and key limiting factors—such as dependency awareness and context fidelity—in realistic, production-grade formal verification settings, thereby providing foundational evidence and practical guidance for LLM-augmented trustworthy software verification.

Assessing LLM performance across different verification projectsEvaluating LLM effectiveness in proof assistant verificationIdentifying factors influencing LLM proof generation success

This work proposes an end-to-end, large language model (LLM)-driven framework that integrates natural language directly into the formal verification pipeline, addressing the longstanding challenge that existing formal verification methods rely on rigorously defined formal specifications and thus struggle to accommodate safety requirements expressed in natural language. The approach leverages an LLM to automatically translate natural language descriptions into formal safety specifications, which are then used to perform compositional verification of code implementations. By circumventing the traditional dependency on manually crafted formal specifications, the method demonstrates a novel pathway toward bridging informal requirements and rigorous verification. Preliminary experiments indicate its feasibility and potential for enhancing code safety, offering a promising direction for making formal verification more accessible and applicable to real-world software development practices.

Code GenerationFormal VerificationLarge Language Models

This work addresses the poor readability of formal proofs, which hinders comprehension by mathematicians. We propose a structure-aware recursive summarization framework that leverages large language models to generate stepwise, informal, and hierarchical natural-language summaries of formal proofs—e.g., those written in Lean. The method recursively compresses subproofs along the proof dependency graph and then integrates contextual information to produce coherent, natural-language explanations. Crucially, it achieves end-to-end generation of highly readable natural-language proofs while strictly preserving logical fidelity. Experiments on textbook-level theorems and the Lean Mathematical Library demonstrate that the generated summaries match or surpass human-written reference proofs in readability, logical faithfulness, and mathematical rigor. These results validate both the method’s effectiveness and its generalizability across diverse mathematical domains.

Evaluating readability and accuracy of generated proofsLeveraging LLMs for informalization and summarizationTranslating formal proofs into natural language

This work proposes a novel paradigm termed “agent-based proof automation” to address the high cost of manually crafting lengthy formal proof scripts. In this approach, human experts supply key mathematical insights, while large language model (LLM) agents autonomously generate and iteratively refine proof scripts within the Lean 4 environment. Relying solely on off-the-shelf LLMs and lightweight verification tools, the method demonstrates— for the first time—the capacity for efficient, large-scale collaborative formal verification by LLM agents. Evaluated on the 14,000-line System Capless type safety proof, the system successfully completed 189 out of 217 tasks (87% success rate), with only 16% of the tasks requiring human intervention.

automationformal verificationlarge language models

Latest Papers

What's happening recently
View more

The 4/$delta$ Bound: Designing Predictable LLM-Verifier Systems for Formal Method Guarantee

Nov 30, 2025
PD
Pierre Dantas
🏛️ The University of Manchester | Federal University of Amazonas

Existing LLM–verifier collaboration frameworks for formal verification lack theoretical guarantees, leading to unstable behavior such as non-termination or divergence. Method: We propose the first formally verified LLM–verifier framework with provable termination and convergence: we model the interaction as a discrete-time Markov chain, establish a quantitative relationship between error-reduction probability δ and expected iteration count, and derive a convergence theorem yielding an analytical upper bound of 4/δ on expected iterations. Contribution/Results: This enables systematic, predictability-driven system design—replacing heuristic tuning with rigorous resource planning. Empirical evaluation across >90,000 tasks demonstrates universal convergence, with measured convergence factor (C_f approx 1.0), confirming tight alignment between theory and practice. The framework provides a quantifiable foundation for resource allocation in safety-critical software verification.

Develops a formal framework with provable guarantees for LLM-verifier convergence and terminationEstablishes design thresholds and predictable performance zones for safety-critical software verificationModels LLM-verifier interaction as a Markov chain using error-reduction probability to bound expected iterations

This work addresses the challenges of ensuring correctness and accurately constructing formal specifications when large language models generate code from natural language. We propose a verifiable code generation approach that integrates hierarchical prompting with verification feedback. To support this, we introduce the NL2VC-60 dataset, which leverages Dafny formal specifications and the uDebug platform to prevent vacuous verification. Furthermore, we design a self-repair prompting mechanism guided by structural signatures and verifier feedback. Experimental results demonstrate that our method substantially enhances both verifiability and functional correctness of code generated by open-source large models: Gemma-4-31B achieves a verification success rate of 90.91%, while GPT-OSS-120B improves from 0% to 81.82% under signature-guided prompting, marking the first systematic validation of open-source models’ potential to produce high-assurance code for complex algorithmic tasks.

Code CorrectnessDafnyFormal Verification

This study addresses the common omission in current large language model evaluations of code generation—the iterative refinement process inherent in real-world programming and the models’ capacity for self-correction using feedback. The authors propose a novel framework that leverages execution-based feedback, such as compilation errors and test failures, to systematically investigate how reasoning and non-reasoning models utilize such signals across multiple programming languages. Through multidimensional categorization of code failures and extensive cross-model, cross-language experiments, they demonstrate that reasoning models consistently improve over iterations and significantly outperform non-reasoning counterparts. While syntactic and runtime errors prove relatively amenable to correction, logical and algorithmic errors remain challenging, thereby delineating the current limits of feedback-driven repair mechanisms.

code correctionexecution feedbackiterative refinement

Large language models often introduce subtle, hard-to-detect bugs when generating complex software, compromising reliability. This work proposes the first fully automated, project-level code generation and verification framework based on an interactive theorem prover (ITP). The approach separates code with side effects into C++ while formalizing pure logical components in the ITP Rocq, where they are automatically verified and extracted for integration. When proofs fail, the concrete counterexample states guide an LLM agent to autonomously repair the code. In experiments, the system generated 1,859 lines of verified Rocq code and extracted 2,848 lines of C++ within 30 minutes, passing 265 unit tests and 12 hours of AFL++ fuzzing with zero crashes or hangs—outperforming Dafny’s backend, which failed to complete verification under identical conditions.

Formal VerificationInteractive Theorem ProvingLarge-scale Code Generation

LLM For Loop Invariant Generation and Fixing: How Far Are We?

Nov 09, 2025
MR
Mostafijur Rahman Akhond
🏛️ York University | Microsoft Research

This work presents the first systematic evaluation of large language models’ (LLMs) ability to infer and repair program loop invariants without auxiliary information. We adopt an empirical framework encompassing diverse open- and closed-source LLMs across multiple scales, integrating domain-knowledge augmentation and few-shot prompting to quantify performance on standard benchmarks for inductive invariant generation and logical defect repair. Results show that LLMs achieve up to 78% success in invariant generation but only 16% in invariant repair—revealing a critical bottleneck in deep logical correction. A key contribution is the identification of auxiliary information—particularly loop semantics prompts and correct examples—as decisive for improving repair accuracy. Our study establishes a reproducible evaluation paradigm for LLM-driven automated program safety analysis and provides concrete, actionable pathways for enhancing invariant repair capabilities.

Assessing performance of various LLMs on inductive loop invariant inferenceEvaluating LLMs' ability to generate and fix loop invariants for programsInvestigating how auxiliary information enhances LLMs' invariant generation and repair

Hot Scholars

LZ

Lihong Zhi

Academy of Mathematics and Systems Science
computer algebraoptimizatoncertified computation
ZY

Zhengfeng Yang

East China Normal University
Symbolic ComputationFormal Methods
CK

Cezary Kaliszyk

Professor, University of Melbourne
Automated ReasoningFormal Methods
JW

Jianlin Wang

Leavey School of Business, Santa Clara University
Financial EconomicsInternational EconomicsMacroeconomics
JW

Jim Woodcock

Professor of Software Engineering, University of York
Software EngineeringFormal Methods