Score
Designs and builds machine-checked formalizations and computer‑aided proofs, including the translation of informal arguments into proof‑assistant scripts, composition and transformation of proofs (equational reasoning, structural induction, constructive and completeness proofs), and the development of proof‑engineering infrastructure for rigorous verification. Implements and analyzes hybrid methods that integrate exhaustive finite‑case enumeration and computational checks with theoretical reductions, and produces verifiable proof traces and proof‑based explanations.
This work systematically integrates and evaluates the full spectrum of advances in artificial intelligence for mathematical reasoning, spanning challenges from informal textual reasoning to formal theorem proving and mathematical discovery. We propose the first unified framework that cohesively combines informal and formal reasoning, multi-agent collaboration, verification loops, and discovery mechanisms. A four-dimensional taxonomy is introduced to critically analyze the fragility of existing approaches. By surveying mainstream techniques—including chain-of-thought prompting, neuro-symbolic systems, autoformalization, and reinforcement learning with verifiable rewards—and their associated benchmarks, we expose critical issues in current evaluations such as saturation, data contamination, and bias. The paper advocates for future research directions centered on verifiable discovery, reasoning efficiency, and accessible infrastructure.
This work proposes a novel paradigm termed “agent-based proof automation” to address the high cost of manually crafting lengthy formal proof scripts. In this approach, human experts supply key mathematical insights, while large language model (LLM) agents autonomously generate and iteratively refine proof scripts within the Lean 4 environment. Relying solely on off-the-shelf LLMs and lightweight verification tools, the method demonstrates— for the first time—the capacity for efficient, large-scale collaborative formal verification by LLM agents. Evaluated on the 14,000-line System Capless type safety proof, the system successfully completed 189 out of 217 tasks (87% success rate), with only 16% of the tasks requiring human intervention.
This work addresses the poor readability, modularity, and maintainability of formal proofs generated by large language models, which often fall short of high-quality mathematical library standards. Inspired by human proof-refactoring practices, the authors propose a four-stage agent framework that systematically decomposes proof refactoring into candidate fragment extraction, auxiliary lemma design, component verification, and original proof repair. Departing from length-based or other single-metric optimizations, the approach prioritizes structural quality. Experiments on Lean-generated proofs from PutnamBench and Putnam2025 demonstrate that the method significantly outperforms the Claude Code baseline in human readability and signature quality, establishing the first automated pipeline for structure-oriented proof refactoring.
This study presents the first systematic evaluation of state-of-the-art large language models’ ability to generate formally verifiable proofs in graduate-level mathematics, including analysis, algebra, probability, and logic. The authors construct a private benchmark that pairs natural language mathematical problems with their formalized statements in Lean 4 and employ an agent-based framework to automatically assess whether model-generated proofs are accepted by the Lean 4 proof checker. Experimental results show that the best base model achieves an accuracy of 33.5%, while other models exhibit substantially lower performance. The work further provides a detailed empirical analysis of tool usage, failure modes, reasoning costs, and latency, offering valuable insights into the current capabilities and limitations of large language models in formal mathematical reasoning.
Automated verification in separation logic (SL) has long relied on ad hoc heuristics, lacking a systematic metatheory and suffering from poor scalability. Method: This paper establishes the first general SL metatheory grounded in category theory and algebraic structures—specifically functors, homomorphisms, and modules over rings—systematically integrating abstract algebra into SL automation. The framework supports compositional model instantiation and modular predicate synthesis for any data structure admitting an algebraic characterization. All results are formally verified in Isabelle/HOL, and an automatic algebraic instantiation algorithm is developed. Contribution/Results: Experiments demonstrate fully automated algebraic modeling of complex imperative program semantics—including lists, trees, and graphs—and yield inference engines whose performance matches state-of-the-art hand-crafted systems. This approach decisively overcomes the scalability limitations inherent in heuristic-based methods.
This study addresses the critical challenge of controlling proof difficulty consistency in the automatic generation of mathematical proof exercises. It proposes a method based on a cut-free semantic tableau proof system within first-order logic, which is free of logical symbols and endowed with analyticity and structural properties. This framework enables the mechanized extraction of rules that capture the cognitive effort required for informal proofs and supports a formal model of proof complexity. Leveraging this approach, the system can generate novel exercises whose difficulty closely matches that of a given problem, making it suitable for discrete mathematics courses. This work represents the first integration of tableau-based structural analysis with formal complexity modeling to achieve controllable, automated generation of proof exercises tailored to educational contexts.
This work addresses the challenge of formally verifying mature, safety-critical industrial C++ codebases by strategically integrating theorem proving (PVS) and model checking (SeaHorn), augmented with large language models to assist in specification construction. The approach is applied to the core order book algorithm of Stellar’s SDEX blockchain module. The verification effort successfully establishes critical correctness properties—including state consistency and unreachability of erroneous states—uncovers discrepancies between documentation and implementation, and produces reusable formal artifacts. These assets enable continuous validation of invariants during future code evolution, thereby enhancing long-term reliability and maintainability of the system.
Existing automated proof synthesis methods struggle with complex theorems in interactive theorem provers and rely heavily on expert knowledge. This work presents the first systematic analysis of failed proof attempts, uncovering critical correlations between human expert proof patterns and successful proofs. Building on these insights, we propose Pattern-Guided Tactic Search (PGTS), a novel approach that integrates deep learning–driven proof synthesis, empirical analysis of proof scripts, and heuristic tactic search guided by expert-derived patterns. Experimental results demonstrate that PGTS improves upon existing tools by proving 8.05% more theorems on standard benchmarks on average and achieves a 20% higher success rate on previously unproven theorems, while also generating more concise proof scripts.
This study addresses the limitations of traditional pen-and-paper instruction in formal proof construction—namely slow iteration cycles, difficulty in error correction, and insufficient student confidence. To overcome these challenges, the authors design and implement an educational, web-based interactive theorem prover that uniquely unifies support for both classical and constructive logics within a single platform, accommodating natural deduction and sequent calculus alike. Built on modern web frontend technologies and integrated with a logical inference engine, the system offers real-time syntax checking, proof state tracking, and dynamic visualization of proof trees. An evaluation involving 35 students demonstrates that the tool significantly enhances learners’ comprehension of formal proofs and engagement with the material; user feedback confirms it effectively accelerates iterative refinement, simplifies debugging, and bolsters problem-solving confidence.
Current automated formalization tools struggle to independently handle the formalization of complex mathematical proofs. This study investigates how human experts conduct proof formalization with AI assistance through a mixed-methods approach, combining qualitative inquiry with controlled user experiments across diverse domains and difficulty levels. It provides the first systematic characterization of how users flexibly orchestrate multiple AI tools in real-world scenarios and reveals a central human need to retain high-level control in human-AI collaboration. The findings demonstrate that AI assistance significantly improves formalization accuracy, and users consistently adapt their tool usage dynamically based on task requirements.