ProofJudge: Tool-Grounded LLM Evaluation of Formal Proof Quality in Mathlib

📅 2026-08-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文介绍ProofJudge系统,通过五个维度评估Lean 4形式证明的质量,并使用工具访问来提高评分准确性。
📝 Abstract
Formal proofs in Lean 4 that pass the kernel's type checker can nonetheless vary widely in quality. We introduce ProofJudge, an agentic LLM-as-judge system that scores formal proof quality along five dimensions beyond correctness: library leverage, automation fit, structural clarity, statement quality, and Mathlib conventions. We evaluate ProofJudge on a novel dataset of 218 declarations drawn from distinct Mathlib PRs. The judge agent is grounded by tool access to the commit the PR is applied to, enabling it to query the library state when scoring. A judge is considered aligned with human preferences when it rates the version of the PR Mathlib accepted above the initial version that was sent back for revision. All six judge models evaluated recover the reviewers' preference well above chance, from 80.8% to 63.5%, and two open-weight judges reach roughly 70% at a tenth of the best judge's cost. We release the judge harness, evaluation dataset, and evaluation traces as open-source artifacts to support further research.
Problem

Research questions and friction points this paper is trying to address.

formal proof quality
Lean 4
type checker
library leverage
automation fit
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-as-judge
formal proof quality
library state access
alignment with human preferences
open-source artifacts
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shane Caldwell
Dreadnode