Score
Designs and implements formal mathematical definitions, theorems, proofs, and library APIs in the Lean theorem prover’s mathlib; integrates new constructions into mathlib by encoding objects and statements, proving lemmas and equivalences, and adapting code to mathlib conventions and automation.
This work addresses the lack of a systematic, formalized knowledge base for computer science in Lean, which has hindered its adoption in education, research, and large-scale verification. To bridge this gap, we present CSLib—the first open-source library of formalized theorems and data structures specifically designed for computer science, built upon the Lean proof assistant and dependent type theory. CSLib establishes a reusable and composable formal infrastructure that significantly expands Lean’s foundational knowledge base in computer science. By providing a comprehensive and extensible collection of verified components, CSLib enables effective collaboration between human developers and AI systems in constructing large-scale formally verified software, thereby advancing the broader application and accessibility of formal methods within the field.
To address the difficulty users face in retrieving theorems from mathlib4 due to unfamiliarity with naming conventions and documentation, this paper introduces LeanSearch—the first semantic search engine tailored for the Lean mathematical library. Methodologically: (1) we construct the first evaluable cross-lingual semantic search benchmark mapping natural-language queries to formal theorems; (2) we propose a joint encoding strategy for theorems and their associated docstrings to build a customized dense semantic index over mathlib4; and (3) we implement an end-to-end embedded retrieval system. Our contributions include establishing the first reproducible, evaluable semantic search infrastructure for mathlib4; deploying a publicly accessible service (leansearch.net); and achieving significant improvements in retrieval accuracy and onboarding experience for novice users—thereby facilitating collaborative formalization within the Lean community.
Formal theorem proving is hindered by the scarcity of high-quality bilingual natural language–Lean 4 data. To address this, we propose the first bidirectional synthetic data construction framework tailored for mathematical theorem proving. Our method employs a large language model–driven iterative generation-and-filtering pipeline, integrating rule-guided filtering, mathematical semantic consistency verification, and proof-search feedback to ensure high-fidelity bidirectional translation between natural language and Lean 4. The resulting dataset introduces 21 newly curated International Mathematical Olympiad (IMO) problems and real-world forum proofs, and we publicly release an open-source dataset of 57K problem–proof pairs (on Hugging Face) alongside full implementation code (on GitHub). Experiments demonstrate substantial improvements in LLM performance across formalization translation, proposition understanding, and proof generation—establishing a foundational data resource for mathematical AI.
Formal verification of the Lean 4 kernel’s correctness remains an open challenge. Method: This paper develops the first fully Lean 4–implemented external type checker, formally specifying its type-theoretic semantics and rigorously proving semantic equivalence between the implementation and the formal semantics. The checker supports end-to-end verification of the entire mathlib library (>1 million lines) and achieves 50%–80% of the performance of the C++ reference implementation. Contribution/Results: It presents the first complete formalization of Lean’s type theory within Lean itself; establishes a provably sound correspondence between kernel primitives and semantic inference rules, thereby providing dual reliability guarantees for kernel evolution; and constitutes a critical step toward a fully self-hosting Lean compiler—significantly enhancing the trustworthiness and maintainability of the theorem prover.
To address the challenge of neural theorem provers failing to sustainably generate correct proofs in fully autonomous mode, this paper proposes a human-in-the-loop formal theorem proving framework for Lean. Methodologically, it enables native execution of large language models (LLMs) within Lean—supporting both local and cloud-based models via a plugin-architected integration—and adopts a human-led, model-assisted paradigm featuring lightweight interactive capabilities: step-wise suggestions, goal completion, and premise selection. Technically, it unifies Lean’s plugin infrastructure, an extensible LLM inference engine (CPU/GPU/cloud-compatible), formal mathematics fine-tuning, and a real-time proof-state interaction interface. Experiments on the *Mathematics in Lean* dataset show that human–AI collaboration requires only 2.08 average manual interventions per proof (outperforming aesop’s 3.86), while achieving a 74.2% fully automated step-wise success rate—a 85% improvement over baseline. All code and models are released under the MIT License.
Existing infrastructure struggles to meet the demands of AI-driven mathematical research for Lean 4, particularly in high-throughput processing, scalable verification, multi-version support, and request-level isolation. This work proposes the first cloud-native Lean 4 service platform, which uniquely enables high concurrency, per-request isolation, and coexistence of multiple Lean 4 and Mathlib versions. The platform integrates 14 metaprogramming tools—including proof checking, semantic source code manipulation, deterministic repair, and lemma extraction—and provides seamless access via HTTP API, Python SDK, CLI, and a web UI, eliminating the need for local deployment. Already publicly deployed, it has processed over 500 million requests and powered Axiom Math’s perfect score in the 2025 Putnam Competition, thereby addressing a critical gap in scalable theorem-proving infrastructure.
为了解决Lean证明助手库向其他系统翻译的问题,本文在Dedukti逻辑框架中提出了一种编码Lean术语和类型的理论,并定义了一个保持可类型化的从Lean大部分子集到该Dedukti理论的翻译方法。
This work addresses the challenge of subtle errors in mathematical reasoning by large language models through a novel multi-agent framework built upon general-purpose code-oriented large language models. The framework employs a coordinator to dynamically orchestrate a customized pipeline for automatically formalizing research-level mathematical theorems in Lean 4. Its key innovation lies in the ability to dynamically extend type definitions and verify auxiliary lemmas without introducing additional axioms. The approach successfully formalizes the core theorems of five STOC papers—two of which rely solely on the Lean kernel—and produces machine-verified proofs for 32 problems on PutnamBench. All formalizations have been expert-reviewed and are publicly released.
本文介绍ProofJudge系统,通过五个维度评估Lean 4形式证明的质量,并使用工具访问来提高评分准确性。
This work addresses the challenge of efficiently determining whether pull requests (PRs) in Mathlib meet the criteria for merging under its manual review process. To this end, it introduces MathlibPR, the first benchmark dataset derived from real-world PR histories in Mathlib4, repurposed as supervised signals. The study proposes a staged evaluation protocol to systematically assess the capability of large language models—including DeepSeek, Qwen, Goedel, and Kimina—and coding agents such as Codex and Claude Code in judging PR merge-readiness. Experimental results reveal that current models struggle to distinguish between PRs that are immediately mergeable and those that pass basic checks yet are ultimately revised or rejected, highlighting the task’s inherent difficulty and laying the groundwork for future development of code review assistance tools and reward models.