Score
Designs and builds machine-checkable formalizations, definitions, proof scripts, tactics, and verification artifacts in the Lean theorem prover (Lean 4), producing mechanized theorems, lemmas, and soundness proofs that can be checked by the system. Analyzes and verifies the correctness and completeness of those formalizations (including elimination of placeholders and generation of machine‑checked certificates) and may use AI-assisted synthesis to translate or generate Lean proofs from informal descriptions.
This paper addresses the high cognitive barrier for secondary-school students and the poor pedagogical fit of existing formal tools in mathematics education. We systematically analyze Lean 4’s architecture—particularly its dependent type system and tactic-based metaprogramming DSL—through formal library evaluation, empirical verification on canonical mathematical theorems, and comparative analysis against Coq and Isabelle. Our study reveals Lean 4’s integrated advantages in proof efficiency, interactive usability, and ecosystem maturity. Crucially, this work presents the first holistic assessment of Lean 4 across three dimensions: automated reasoning capability, runtime performance, and pedagogical accessibility. We thereby establish Lean 4’s dual potential as a foundational infrastructure for both mathematics education and lightweight industrial verification. Our findings provide theoretical grounding and actionable pathways for scaling formal methods in secondary mathematics curricula and resource-constrained verification settings. (149 words)
Formal verification of the Lean 4 kernel’s correctness remains an open challenge. Method: This paper develops the first fully Lean 4–implemented external type checker, formally specifying its type-theoretic semantics and rigorously proving semantic equivalence between the implementation and the formal semantics. The checker supports end-to-end verification of the entire mathlib library (>1 million lines) and achieves 50%–80% of the performance of the C++ reference implementation. Contribution/Results: It presents the first complete formalization of Lean’s type theory within Lean itself; establishes a provably sound correspondence between kernel primitives and semantic inference rules, thereby providing dual reliability guarantees for kernel evolution; and constitutes a critical step toward a fully self-hosting Lean compiler—significantly enhancing the trustworthiness and maintainability of the theorem prover.
This work addresses the challenge of integrating dependent type theory in Lean 4 with automated theorem provers (ATPs). We propose the first formally verified, reliable translation from dependent types to first-order logic (FOL). Methodologically, we leverage Lean 4’s metaprogramming capabilities to perform dependent type erasure and structured FOL encoding, and integrate SMT/ATP tools—including Z3 and Vampire—via a novel, general-purpose ATP interface. Our key contributions are: (i) the first formal verification of translation correctness within a dependent type system; and (ii) significant improvements in automated proof success rates on real mathematical lemmas from the Mathlib4 benchmark, surpassing prior tools’ limitations. This work establishes a new paradigm for interactive theorem provers that jointly achieves high expressive power and robust automation.
This work addresses critical reliability issues in existing Lean theorem-proving benchmarks, where inconsistencies between formal statements and informal problem descriptions, along with susceptibility to trivial or adversarial solutions, undermine evaluation validity. To tackle this, the study introduces the first fault taxonomy for formal mathematical datasets and develops an automated auditing toolkit integrating static program analysis, formal verification, semantic auditing via prompt engineering, and manual validation. Applying this framework, the authors conduct a large-scale audit of five prominent benchmarks, uncovering 4,833 issues—including 398 severe defects—and demonstrate that uncorrected flaws significantly distort prover rankings. The paper releases both the auditing tools and corrected dataset snapshots to foster reproducible and trustworthy evaluations in theorem proving.
To address the challenge of neural theorem provers failing to sustainably generate correct proofs in fully autonomous mode, this paper proposes a human-in-the-loop formal theorem proving framework for Lean. Methodologically, it enables native execution of large language models (LLMs) within Lean—supporting both local and cloud-based models via a plugin-architected integration—and adopts a human-led, model-assisted paradigm featuring lightweight interactive capabilities: step-wise suggestions, goal completion, and premise selection. Technically, it unifies Lean’s plugin infrastructure, an extensible LLM inference engine (CPU/GPU/cloud-compatible), formal mathematics fine-tuning, and a real-time proof-state interaction interface. Experiments on the *Mathematics in Lean* dataset show that human–AI collaboration requires only 2.08 average manual interventions per proof (outperforming aesop’s 3.86), while achieving a 74.2% fully automated step-wise success rate—a 85% improvement over baseline. All code and models are released under the MIT License.
Existing infrastructure struggles to meet the demands of AI-driven mathematical research for Lean 4, particularly in high-throughput processing, scalable verification, multi-version support, and request-level isolation. This work proposes the first cloud-native Lean 4 service platform, which uniquely enables high concurrency, per-request isolation, and coexistence of multiple Lean 4 and Mathlib versions. The platform integrates 14 metaprogramming tools—including proof checking, semantic source code manipulation, deterministic repair, and lemma extraction—and provides seamless access via HTTP API, Python SDK, CLI, and a web UI, eliminating the need for local deployment. Already publicly deployed, it has processed over 500 million requests and powered Axiom Math’s perfect score in the 2025 Putnam Competition, thereby addressing a critical gap in scalable theorem-proving infrastructure.
This study investigates the reasoning capabilities and limitations of artificial intelligence in tackling formalized mathematical problems, exemplified by the “grasshopper problem” (IMO 2009, Problem 6). Using the Aristotle API to generate Lean 4 proofs, the authors construct a formal framework comprising four verified auxiliary lemmas, employing techniques such as maximality arguments, adjacent-swap strategies, and partial-sum properties. While local reasoning succeeds in establishing these lemmas, the main theorem remains incomplete due to an unresolved global counting contradiction, marked with a ‘sorry’ placeholder—highlighting a critical bottleneck in AI’s ability to conduct combinatorial global reasoning. This work presents the first verifiable, AI-generated formal proof artifact that clearly delineates proven from unproven components and releases reproducible Lean code, offering a precise diagnostic of current limitations in AI-assisted theorem proving.
This work addresses the challenge of achieving efficient and reliable formal verification of production-grade cryptographic code written in Rust. We present the first end-to-end Rust-to-Lean 4 verification pipeline, integrating the Charon, Aeneas, and Hax frameworks for symbolic extraction, leveraging the ArkLib and CompPoly libraries of formally specified cryptographic primitives, and introducing the Aristotle and Aleph AI-powered provers to automatically discharge complex proof obligations. All results are rigorously validated by the Lean 4 kernel. Our approach successfully reproduces and fully verifies key cryptographic primitives from Plonky3 and RISC Zero—including FRI folding, finite field arithmetic, Horner evaluation, and Merkle inclusion proofs—and automatically completes proofs for two longstanding open conjectures.
This work presents the first complete and fully verified machine-checked proof of the completeness theorem ($\Gamma \models \varphi \Rightarrow \Gamma \vdash \varphi$) for hybrid logic $L(\forall)$ in Lean 4, eliminating the use of unverified axioms ("sorry"). Building on Oltean’s formalization, we introduce two structural mechanisms for handling fresh names: a namespace-preserving strategy for constructing root witness sets and a Henkin-style construction augmented with data accumulators to manage diamond-modality successors. By integrating Lindenbaum extensions, the Henkin existence lemma, compactness arguments, and a carefully designed axiomatic system—and relying solely on the three foundational axioms `propext`, `Classical.choice`, and `Quot.sound`—we achieve a rigorous and executable completeness proof that resolves the longstanding gap in prior formalizations.
Current approaches to generating Lean formal statements from natural language rely solely on compilation success rates, which often fail to ensure semantic faithfulness, leading to issues such as omitted premises, incorrect domain specifications, or trivial propositions. This work proposes a novel evaluation paradigm centered on “consensus faithfulness,” introducing a benchmark dataset of 400 graduate-level mathematical problems. Through an integrated methodology combining Lean compilation checks, multi-model semantic evaluations, expert calibration, and a 2³ factorial experiment, the study systematically identifies key factors affecting faithfulness. Experiments reveal that even the best tool-augmented agent achieves only 60.5% consensus faithfulness despite an 89.5% compilation rate, and manual auditing confirms this metric’s high conservativeness, highlighting a substantial gap between syntactic compilability and genuine semantic accuracy.