Score
Designs, builds, and audits machine-checkable formal proofs and verification artifacts — including mechanized theorems, formal soundness and safety proofs, certified bounds, and encoded guards/verification conditions — using proof assistants (e.g., Lean 4), SMT solvers, and automated theorem provers. Engineers the proof pipeline and tool integration (compiler emission of proofs, translation of conditions to SMT or proof assistant queries, proof obligation encoding), and analyzes and classifies proof failures or open obligations to verify cross-system mappings and maintain proof-checking correctness.
This work addresses the critical challenge of reliably integrating automated reasoning tools—such as theorem provers, SAT/SMT solvers, and termination analyzers—with proof assistants to build highly trustworthy systems. It presents a systematic survey and comparative analysis of two principal technical approaches: certification and formal verification. The study examines core methodologies including logical encoding, result replay and checking, and integration mechanisms within proof assistants. By elucidating the respective strengths and limitations of these methods and illustrating them through multiple successful case studies, the paper offers clear methodological guidance for constructing high-assurance automated reasoning systems, thereby substantially enhancing the verifiability and trustworthiness of their outputs.
Lean lacks SMT-driven automated proof capabilities comparable to Isabelle/HOL’s Sledgehammer. This paper presents the first end-to-end solution in Lean for generating and faithfully reconstructing SMT proofs: it automatically encodes Lean goals into SMT-LIB, invokes external solvers (e.g., Z3, CVC5) for verification, and reliably reconstructs their proofs as checkable, native Lean terms. The approach leverages Lean’s metaprogramming framework and a custom reconstruction algorithm, significantly reducing the trusted computing base while preserving logical soundness and enhancing automation. Evaluated on the Sledgehammer benchmark suite, it achieves strong performance. As a standalone SMT-LIB proof checker, it attains high verification success rates, operates with a minimal trusted base, and incurs only moderate runtime overhead.
This paper addresses the high cognitive barrier for secondary-school students and the poor pedagogical fit of existing formal tools in mathematics education. We systematically analyze Lean 4’s architecture—particularly its dependent type system and tactic-based metaprogramming DSL—through formal library evaluation, empirical verification on canonical mathematical theorems, and comparative analysis against Coq and Isabelle. Our study reveals Lean 4’s integrated advantages in proof efficiency, interactive usability, and ecosystem maturity. Crucially, this work presents the first holistic assessment of Lean 4 across three dimensions: automated reasoning capability, runtime performance, and pedagogical accessibility. We thereby establish Lean 4’s dual potential as a foundational infrastructure for both mathematics education and lightweight industrial verification. Our findings provide theoretical grounding and actionable pathways for scaling formal methods in secondary mathematics curricula and resource-constrained verification settings. (149 words)
Existing code-level formal verification tools scale poorly to large-scale software, while mainstream unit-level verification relies heavily on manual effort, often missing critical defects. This paper proposes the “Unit Proof Framework” research agenda—the first systematic definition of a unit verification paradigm supporting automated decoupling and independent verification of code units. Methodologically, it integrates formal verification, program analysis, modular verification, and automated toolchain design, with deep alignment to industrial development practices (e.g., AWS workflows). Its core contributions include: (1) establishing a scalable, engineering-friendly unit verification methodology; (2) characterizing a taxonomy of key technical challenges; (3) overcoming bottlenecks inherent in manual verification; and (4) significantly improving early detection of code-level defects. Collectively, this work lays the theoretical foundation and provides a practical technical pathway for building high-assurance, deployable automated verification infrastructure.
This work proposes a novel paradigm termed “agent-based proof automation” to address the high cost of manually crafting lengthy formal proof scripts. In this approach, human experts supply key mathematical insights, while large language model (LLM) agents autonomously generate and iteratively refine proof scripts within the Lean 4 environment. Relying solely on off-the-shelf LLMs and lightweight verification tools, the method demonstrates— for the first time—the capacity for efficient, large-scale collaborative formal verification by LLM agents. Evaluated on the 14,000-line System Capless type safety proof, the system successfully completed 189 out of 217 tasks (87% success rate), with only 16% of the tasks requiring human intervention.
This work addresses a central challenge in system security: formally verifying that system designs and implementations satisfy intended safety properties and support security certification. The authors propose a systematic approach grounded in proof assistants, integrating interactive theorem proving and formal methods to precisely model and machine-check critical security properties across diverse domains—including system security, language-level security, secure compilation, and cryptography. By enabling rigorous, machine-verifiable proofs of correctness, this methodology significantly strengthens the formal assurance of security properties and provides a unified theoretical framework and toolchain for constructing verifiable and certifiable secure systems.
This work addresses the challenge of formally verifying mature, safety-critical industrial C++ codebases by strategically integrating theorem proving (PVS) and model checking (SeaHorn), augmented with large language models to assist in specification construction. The approach is applied to the core order book algorithm of Stellar’s SDEX blockchain module. The verification effort successfully establishes critical correctness properties—including state consistency and unreachability of erroneous states—uncovers discrepancies between documentation and implementation, and produces reusable formal artifacts. These assets enable continuous validation of invariants during future code evolution, thereby enhancing long-term reliability and maintainability of the system.
This work addresses critical reliability issues in existing Lean theorem-proving benchmarks, where inconsistencies between formal statements and informal problem descriptions, along with susceptibility to trivial or adversarial solutions, undermine evaluation validity. To tackle this, the study introduces the first fault taxonomy for formal mathematical datasets and develops an automated auditing toolkit integrating static program analysis, formal verification, semantic auditing via prompt engineering, and manual validation. Applying this framework, the authors conduct a large-scale audit of five prominent benchmarks, uncovering 4,833 issues—including 398 severe defects—and demonstrate that uncorrected flaws significantly distort prover rankings. The paper releases both the auditing tools and corrected dataset snapshots to foster reproducible and trustworthy evaluations in theorem proving.
This work addresses the poor readability, modularity, and maintainability of formal proofs generated by large language models, which often fall short of high-quality mathematical library standards. Inspired by human proof-refactoring practices, the authors propose a four-stage agent framework that systematically decomposes proof refactoring into candidate fragment extraction, auxiliary lemma design, component verification, and original proof repair. Departing from length-based or other single-metric optimizations, the approach prioritizes structural quality. Experiments on Lean-generated proofs from PutnamBench and Putnam2025 demonstrate that the method significantly outperforms the Claude Code baseline in human readability and signature quality, establishing the first automated pipeline for structure-oriented proof refactoring.
This work proposes a large language model–driven automated theorem proving system that enables human–machine collaborative formal verification. The system employs a Planner–Worker–Verifier multi-agent architecture to decompose proof tasks into parallel subgoals, integrates Lean 4 for automatic formal verification, and manages intermediate reasoning through a shared whiteboard and knowledge base. Innovatively combining agent-based automated proving with interactive user guidance within an open-source framework, it provides a terminal interface to support reproducible collaborative exploration. Experimental results on the ProofNet benchmark demonstrate that the approach significantly outperforms simple baselines. The system is fully open-sourced and designed for reproducible evaluation.