Score
Design and implement metaprograms and tactics for the Lean theorem prover that inspect and transform goals, proofs, and the environment to automate routine proof steps and protocol interactions. Build and extend Lean libraries by integrating protocol-specific automation and exposing ergonomic library interfaces for composing and reusing automation.
This paper addresses the high cognitive barrier for secondary-school students and the poor pedagogical fit of existing formal tools in mathematics education. We systematically analyze Lean 4’s architecture—particularly its dependent type system and tactic-based metaprogramming DSL—through formal library evaluation, empirical verification on canonical mathematical theorems, and comparative analysis against Coq and Isabelle. Our study reveals Lean 4’s integrated advantages in proof efficiency, interactive usability, and ecosystem maturity. Crucially, this work presents the first holistic assessment of Lean 4 across three dimensions: automated reasoning capability, runtime performance, and pedagogical accessibility. We thereby establish Lean 4’s dual potential as a foundational infrastructure for both mathematics education and lightweight industrial verification. Our findings provide theoretical grounding and actionable pathways for scaling formal methods in secondary mathematics curricula and resource-constrained verification settings. (149 words)
This work addresses the challenge of integrating dependent type theory in Lean 4 with automated theorem provers (ATPs). We propose the first formally verified, reliable translation from dependent types to first-order logic (FOL). Methodologically, we leverage Lean 4’s metaprogramming capabilities to perform dependent type erasure and structured FOL encoding, and integrate SMT/ATP tools—including Z3 and Vampire—via a novel, general-purpose ATP interface. Our key contributions are: (i) the first formal verification of translation correctness within a dependent type system; and (ii) significant improvements in automated proof success rates on real mathematical lemmas from the Mathlib4 benchmark, surpassing prior tools’ limitations. This work establishes a new paradigm for interactive theorem provers that jointly achieves high expressive power and robust automation.
Lean lacks SMT-driven automated proof capabilities comparable to Isabelle/HOL’s Sledgehammer. This paper presents the first end-to-end solution in Lean for generating and faithfully reconstructing SMT proofs: it automatically encodes Lean goals into SMT-LIB, invokes external solvers (e.g., Z3, CVC5) for verification, and reliably reconstructs their proofs as checkable, native Lean terms. The approach leverages Lean’s metaprogramming framework and a custom reconstruction algorithm, significantly reducing the trusted computing base while preserving logical soundness and enhancing automation. Evaluated on the Sledgehammer benchmark suite, it achieves strong performance. As a standalone SMT-LIB proof checker, it attains high verification success rates, operates with a minimal trusted base, and incurs only moderate runtime overhead.
To address the challenge of neural theorem provers failing to sustainably generate correct proofs in fully autonomous mode, this paper proposes a human-in-the-loop formal theorem proving framework for Lean. Methodologically, it enables native execution of large language models (LLMs) within Lean—supporting both local and cloud-based models via a plugin-architected integration—and adopts a human-led, model-assisted paradigm featuring lightweight interactive capabilities: step-wise suggestions, goal completion, and premise selection. Technically, it unifies Lean’s plugin infrastructure, an extensible LLM inference engine (CPU/GPU/cloud-compatible), formal mathematics fine-tuning, and a real-time proof-state interaction interface. Experiments on the *Mathematics in Lean* dataset show that human–AI collaboration requires only 2.08 average manual interventions per proof (outperforming aesop’s 3.86), while achieving a 74.2% fully automated step-wise success rate—a 85% improvement over baseline. All code and models are released under the MIT License.
Formal verification of the Lean 4 kernel’s correctness remains an open challenge. Method: This paper develops the first fully Lean 4–implemented external type checker, formally specifying its type-theoretic semantics and rigorously proving semantic equivalence between the implementation and the formal semantics. The checker supports end-to-end verification of the entire mathlib library (>1 million lines) and achieves 50%–80% of the performance of the C++ reference implementation. Contribution/Results: It presents the first complete formalization of Lean’s type theory within Lean itself; establishes a provably sound correspondence between kernel primitives and semantic inference rules, thereby providing dual reliability guarantees for kernel evolution; and constitutes a critical step toward a fully self-hosting Lean compiler—significantly enhancing the trustworthiness and maintainability of the theorem prover.
This work proposes a large language model–driven automated theorem proving system that enables human–machine collaborative formal verification. The system employs a Planner–Worker–Verifier multi-agent architecture to decompose proof tasks into parallel subgoals, integrates Lean 4 for automatic formal verification, and manages intermediate reasoning through a shared whiteboard and knowledge base. Innovatively combining agent-based automated proving with interactive user guidance within an open-source framework, it provides a terminal interface to support reproducible collaborative exploration. Experimental results on the ProofNet benchmark demonstrate that the approach significantly outperforms simple baselines. The system is fully open-sourced and designed for reproducible evaluation.
Existing infrastructure struggles to meet the demands of AI-driven mathematical research for Lean 4, particularly in high-throughput processing, scalable verification, multi-version support, and request-level isolation. This work proposes the first cloud-native Lean 4 service platform, which uniquely enables high concurrency, per-request isolation, and coexistence of multiple Lean 4 and Mathlib versions. The platform integrates 14 metaprogramming tools—including proof checking, semantic source code manipulation, deterministic repair, and lemma extraction—and provides seamless access via HTTP API, Python SDK, CLI, and a web UI, eliminating the need for local deployment. Already publicly deployed, it has processed over 500 million requests and powered Axiom Math’s perfect score in the 2025 Putnam Competition, thereby addressing a critical gap in scalable theorem-proving infrastructure.
This work addresses the challenge of integrating the industrial-scale B-Method tool Atelier B with the Lean proof assistant by introducing BARReL, a library implemented in Lean 4 that enables users to carry out the entire development process—from formal specification to machine refinement—using standard B syntax within Lean. The key innovation lies in leveraging Lean’s dependent type system to explicitly encode well-definedness conditions for partial B operators, thereby ensuring that all proof obligations are free from ill-formed instances. Furthermore, metaprogramming is employed to automatically generate well-definedness constraints and basic automation tactics. The approach has been validated on representative case studies, laying the foundation for a highly reliable and extensible, Lean-native toolchain for the B Method.
Current approaches to generating Lean formal statements from natural language rely solely on compilation success rates, which often fail to ensure semantic faithfulness, leading to issues such as omitted premises, incorrect domain specifications, or trivial propositions. This work proposes a novel evaluation paradigm centered on “consensus faithfulness,” introducing a benchmark dataset of 400 graduate-level mathematical problems. Through an integrated methodology combining Lean compilation checks, multi-model semantic evaluations, expert calibration, and a 2³ factorial experiment, the study systematically identifies key factors affecting faithfulness. Experiments reveal that even the best tool-augmented agent achieves only 60.5% consensus faithfulness despite an 89.5% compilation rate, and manual auditing confirms this metric’s high conservativeness, highlighting a substantial gap between syntactic compilability and genuine semantic accuracy.
This work proposes a self-evolving Lean theorem-proving agent designed to autonomously optimize its proof workflow without reliance on manually engineered strategies. The approach introduces a verifier-anchored co-evolution mechanism that leverages dynamic curriculum learning and single-anchor recalibration, enabling the agent to co-evolve alongside its benchmark while preserving score comparability and ensuring progressive difficulty. The system is built upon a Lean-based verification loop, a self-modifying architecture, and a machine-readable representation of proof context. Evaluated on the miniF2F test set, the co-evolving agent achieves a success rate of 45.1%, substantially outperforming both a fixed-benchmark agent (32.0%) and the initial seed model (12.7%), thereby demonstrating the efficacy and novelty of the proposed method.