interactive theorem proving

Designs, builds, and evaluates interactive systems that support stepwise, user-driven construction and formal verification of proofs—including theorem provers, educational or web-based interfaces—and the engines that perform incremental or real-time proof checking. These artifacts enforce correct rule application and provide immediate, live feedback to validate entered proof steps and speed iteration and error correction.

interactivetheoremproving

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$228K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Theorem Provers: One Size Fits All?

Sep 18, 2025
HO
Harrison Oates
🏛️ The Australian National University

This study addresses the lack of empirical evidence in theorem prover selection by conducting the first systematic, cross-platform comparison of Coq and Idris2—evaluated on a unified task: correctness verification of insertion sort. The methodology employs interactive formal verification, integrating implementation, proof strategy design, and standard library usage to enable both qualitative and empirical analysis across three dimensions: usability, community support, and library ecosystem. Results indicate that Coq exhibits significant advantages in standard library completeness, toolchain maturity, and community resources. In contrast, Idris2 demonstrates innovative potential in proof expressiveness and program-proof integration, leveraging its dependent type system and built-in computational capabilities. This work establishes the first empirically grounded, task-aligned benchmark for cross-prover evaluation and provides practitioners with actionable guidance for formal tool selection and system design.

Comparing community and library support for proversEvaluating usability of theorem provers Coq and Idris2Guiding informed system choice for formal verification

OnlineProver: Experience with a Visualisation Tool for Teaching Formal Proofs

May 07, 2025
JP
Ján Perháˇc
🏛️ Technical University of Košice | University of Oslo | University of Copenhagen

To address students’ difficulties in comprehending formal proofs and the delayed, non-contextual feedback prevalent in formal methods education, this paper designs and implements OnlineProver—a pedagogy-oriented, web-based interactive proof assistant. Built upon a lightweight, visual natural deduction framework, OnlineProver supports handwritten-style proof construction and delivers real-time, context-aware error feedback, balancing learner autonomy with instructional efficacy. The system adopts a web-service architecture integrating dynamic frontend rendering with a rule-driven proof-checking engine. Its pedagogical impact was empirically validated through classroom deployment and student surveys. Results indicate high user satisfaction and statistically significant improvement in students’ formal reasoning abilities. Crucially, OnlineProver introduces the first “education-first” lightweight visualization paradigm for interactive proof assistants—offering both a reusable design methodology and empirical evidence for pedagogically grounded tool development in formal verification education.

Develop interactive proof assistant for teaching formal proofsEvaluate effectiveness in classroom learning settingsProvide user-friendly interface with real-time feedback

StepFun-Prover Preview: Let's Think and Verify Step by Step

Jul 27, 2025
SS
Shijie Shang
🏛️ StepFun | University of Chinese Academy of Sciences

This work addresses the challenge of unreliable Lean 4 proof generation by large language models (LLMs) in formal theorem proving. We propose a tool-augmented, end-to-end reasoning framework that tightly integrates LLMs with the Lean 4 proof environment, enabling interactive tool invocation, real-time feedback-driven reinforcement learning, and human-like stepwise verification. Our key contribution is a closed-loop “generate–execute–feedback–revise” mechanism, which significantly improves proof-generation reliability under few-shot conditions. Evaluated on the miniF2F-test benchmark, our approach achieves a 70.0% pass@1 success rate—constituting a substantial improvement over prior methods. This framework establishes a new paradigm for automated theorem proving and trustworthy mathematical AI assistants.

Develops a model for formal theorem proving using toolsEnhances proof generation with reinforcement learningImproves automated theorem proving benchmark performance

This work proposes a novel paradigm termed “agent-based proof automation” to address the high cost of manually crafting lengthy formal proof scripts. In this approach, human experts supply key mathematical insights, while large language model (LLM) agents autonomously generate and iteratively refine proof scripts within the Lean 4 environment. Relying solely on off-the-shelf LLMs and lightweight verification tools, the method demonstrates— for the first time—the capacity for efficient, large-scale collaborative formal verification by LLM agents. Evaluated on the 14,000-line System Capless type safety proof, the system successfully completed 189 out of 217 tasks (87% success rate), with only 16% of the tasks requiring human intervention.

automationformal verificationlarge language models

Cobblestone: Iterative Automation for Formal Verification

Oct 25, 2024
SR
Saketh Ram Kasibatla
🏛️ UC San Diego | University of Illinois, Urbana-Champaign | University of Massachusetts, Amherst

To address the high proof-writing cost and low automation in Coq formal verification, this paper proposes an LLM-driven proof generation method based on iterative synthesis. It leverages large language models to batch-generate candidate proofs and innovatively identifies and cross-proofs fuses locally valid fragments from multiple failed attempts. The approach supports incremental synthesis guided by partial progress—including subgoal decomposition and external lemmas. Under a strict no-training-data-leakage constraint, it achieves a 48% fully automated proof rate—31 percentage points higher than the prior SOTA Proverbot9001—and reaches 58% when incorporating external progress, establishing a new end-to-end SOTA for Coq automated theorem proving. Its core innovation lies in the systematic mining and cross-proof recomposition of effective fragments from failed proofs.

Automates formal verification using iterative proof synthesisCombines LLMs with divide-and-conquer to handle complex proofsImproves success rates over existing tools with low cost

Latest Papers

What's happening recently
View more

This study addresses the limitations of traditional pen-and-paper instruction in formal proof construction—namely slow iteration cycles, difficulty in error correction, and insufficient student confidence. To overcome these challenges, the authors design and implement an educational, web-based interactive theorem prover that uniquely unifies support for both classical and constructive logics within a single platform, accommodating natural deduction and sequent calculus alike. Built on modern web frontend technologies and integrated with a logical inference engine, the system offers real-time syntax checking, proof state tracking, and dynamic visualization of proof trees. An evaluation involving 35 students demonstrates that the tool significantly enhances learners’ comprehension of formal proofs and engagement with the material; user feedback confirms it effectively accelerates iterative refinement, simplifies debugging, and bolsters problem-solving confidence.

formal prooflogic educationnatural deduction

This work proposes a large language model–driven automated theorem proving system that enables human–machine collaborative formal verification. The system employs a Planner–Worker–Verifier multi-agent architecture to decompose proof tasks into parallel subgoals, integrates Lean 4 for automatic formal verification, and manages intermediate reasoning through a shared whiteboard and knowledge base. Innovatively combining agent-based automated proving with interactive user guidance within an open-source framework, it provides a terminal interface to support reproducible collaborative exploration. Experimental results on the ProofNet benchmark demonstrate that the approach significantly outperforms simple baselines. The system is fully open-sourced and designed for reproducible evaluation.

automated theorem provingformal verificationinteractive proof

This work addresses the susceptibility of large language models to “context contamination” when verifying complex mathematical proofs, which often masks logical errors. The authors propose a step-level verification framework that preserves the full context of each inference step and strictly restricts the set of admissible theorems for reference, enabling fine-grained validation of research-level proofs. By integrating LLM-guided reasoning constraints, an adversarial benchmark (FirstProof Challenge), and systematic ablation studies, the method substantially outperforms conventional global evaluation approaches in precisely identifying subtle logical flaws. Remaining misjudgments predominantly stem not from severe hallucinations but from “over-rigor”—arising when domain-specific conventions are left implicit—thereby exposing latent ambiguities in existing expert benchmarks and advancing a more human-like, cautious paradigm for mathematical verification.

context poisoninglarge language modelslogical errors

While current large language models can automatically fill proof holes (i.e., eliminate 'sorries') in interactive theorem proving, their generated formalizations often fail expert review due to ill-conceived definitions, insufficiently general theorems, or suboptimal API design. This work presents a semi-autonomous formalization of Grothendieck’s vanishing theorem as a case study and introduces expert review as a central criterion for evaluating the quality of automated formalizations. By integrating large language model assistance, interactive proving, and an iterative refactoring-compression pipeline, the study systematically assesses the high-level design usability of automatically generated content. The findings reveal that measuring success solely by 'sorry' closure is markedly inadequate; expert-driven refactoring substantially improves formalization quality, underscoring the critical role of expert acceptability in evaluating automated formalization efforts.

autoformalizationexpert reviewformalization quality

Program verification is inherently undecidable, and existing tools either lack sufficient user interaction capabilities or operate at an abstraction level too low to enable users to effectively comprehend proof states and correct errors. This work proposes a novel interactive verification approach that, for the first time, enables direct visualization of and intervention in proof states at the source code and specification levels, thereby bridging the cognitive gap between high-level semantics and low-level logical reasoning. A prototype system built upon the Java verification engine KeY integrates automated and interactive techniques to allow users to guide proof exploration at a high semantic level. User studies demonstrate that this method significantly enhances users’ understanding of the verification process and enables more efficient identification of flaws in either code or specifications.

autoactive verificationprogram verificationproof state

Hot Scholars

JA

Jeremy Avigad

Professor of Philosophy and Mathematical Sciences, Carnegie Mellon University
Mathematical logicproof theoryphilosophy of mathematicsformal verification
ZW

Zaiwen Wen

Peking University
OptimizationMachine Learning
LB

Lars Birkedal

Dept. of Computer Science, Aarhus University
Computer ScienceProgrammingLogicSemantics
AP

André Platzer

Alexander von Humboldt Professor, Karlsruhe Institute of Technology
Formal MethodsLogic in Computer ScienceTheorem ProvingProgramming Languages
DM

Dale Miller

Inria-Saclay and LIX, Ecole Polytechnique
Proof TheoryLinear LogicLogic ProgrammingTheorem Proving