AIProver: Agentic Auto-Formalization of Mathematical Research via Certificate-Driven Evolving Harness

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of missing libraries, semantic distortion, and data scarcity in automated mathematical proof formalization by proposing HarnessEvolve, a certificate-driven mechanism that achieves autonomous proof synthesis through joint post-training of open-source large language models within an evolutionary agent framework. The method integrates symbolic feedback reinforcement learning with type and semantic verifiers to co-optimize model behavior and control flow, alongside the construction of the LoCoBench benchmark dataset. Experimental results demonstrate that the proposed system significantly outperforms existing state-of-the-art approaches, achieving substantially higher semantic correctness than closed-source tools such as Claude while reducing computational costs by 24%. This work effectively overcomes the reliance of automated formalization on commercial closed-source models.
📝 Abstract
Proof auto-formalization translates natural-language (NL) theorems and proofs into a formal language (FL) such as Lean, enabling mechanical verification. Despite rapid progress, research-level proofs often depend on concepts missing from leading proof assistant libraries (e.g., Lean's Mathlib), and successful compilation does not guarantee that a translation preserves the theorem's meaning or the proof's reasoning. Furthermore, aligned NL-FL training data are scarce, and leading agents often rely on costly frontier models and manually engineered harnesses. To address the above issues, we present AIProver, an agentic framework for autonomous proof auto-formalization and proof synthesis (AFPS) that jointly post-trains a 119B open-weight language model and evolves its agentic, tool-calling harness with HarnessEvolve. Verifiers assess type correctness, proof completeness, and semantic correctness, returning rewards and diagnostic certificates that drive model fine-tuning and alternating reinforcement learning via symbolic feedback and HarnessEvolve, a certificate-driven evolutionary search over the whole harness control flow that re-tailors the harness to the updated model. For research-level training and evaluation, we introduce LoCoBench, 58.9k instances from Mathlib, CSLib, Mizar Math Library, and a bounded-arithmetic textbook, with a 771-instance validation split whose theorem-proof pairs have no public Lean formalization. Against 39 frameworks spanning AFPS agents, frontier LLMs, and coding agents, AIProver lifts pass@4 semantic correctness over its Leanstral-1.5 base from 15.7% to 36.7% and outperforms every other open-weight system and Aristotle. As a Claude Code and Codex skill, it lifts their semantic correctness from 41.9% and 34.1% to 79.8% and 62.4%, respectively. Further, it is also 24% cheaper than Numina-Lean-Agent, pushing the accuracy-cost frontier of research-level AFPS.
Problem

Research questions and friction points this paper is trying to address.

proof auto-formalization
semantic correctness
aligned training data
research-level proofs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Auto-formalization
Agentic framework
Certificate-driven evolution
Reinforcement learning
Proof synthesis
🔎 Similar Papers
No similar papers found.