ProofGap: Benchmarking Step-Level Formal Reasoning with Local Obligations Derived from Natural-Language Solutions

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing formal reasoning benchmarks, which evaluate only complete proofs and lack fine-grained, step-level diagnostic capabilities. We propose a novel method that generates localized proof obligations from natural language solutions. By leveraging a natural language processing pipeline and semantic alignment techniques, mathematical analysis exercises are decomposed into context- and goal-isolated proof gaps. Based on the Demidovich problem set, we construct a fine-grained benchmark dataset comprising 3,015 exercises and 26,116 proof gaps, achieving a breakthrough in evaluation granularity from the theorem level to the step level. This benchmark precisely localizes model failure points during formal proof construction, significantly enhancing diagnostic precision and providing foundational support for the future development of proof verification systems.
📝 Abstract
Existing formal mathematics benchmarks, such as miniF2F, ProofNet, and PutnamBench, primarily evaluate models on constructing complete formal proofs for challenging problems. Because success is measured at the theorem level, these benchmarks offer limited insight into models' step-level formal reasoning. Evaluating this capability separately enables finer-grained diagnosis of model limitations than theorem-level evaluation alone. To fill this evaluation gap, we introduce ProofGap, a fine-grained benchmark for step-level formal reasoning. ProofGap is constructed through a natural-language proof-processing pipeline that decomposes each reasoning step into one or more aligned proof gaps. Applying this pipeline to natural-language solutions to 3,015 exercises in B. P. Demidovich's Problems in Mathematical Analysis yields 26,116 gaps. The benchmark focuses on mathematical analysis, a domain that remains challenging for current models. By supplying the local context and target explicitly, gap completion isolates local formal proof construction from end-to-end proof composition, enabling more precise localization of model failures. Natural-language solutions serve as the provenance of these obligations, while the benchmark task itself starts from an already formalized local context and goal. Beyond benchmarking, the same pipeline may support future proof-verification systems, provided that semantic translation and sequential proof composition are handled reliably.
Problem

Research questions and friction points this paper is trying to address.

formal reasoning
step-level evaluation
benchmarking
proof gaps
mathematical analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Step-Level Formal Reasoning
ProofGap Benchmark
Natural-Language Proof Processing
Local Obligations
Mathematical Analysis
L
Lihan Xie
Shanghai Jiao Tong University; Shanghai Innovation Institute
Z
Zhicheng Hui
Shanghai Jiao Tong University
Y
Yingjun Lan
Shanghai Jiao Tong University
Z
Zhehao Li
Shanghai Jiao Tong University
X
Xingzhi Qi
Shanghai Jiao Tong University
S
Siyue Huang
Shanghai Jiao Tong University
J
Jirui Liu
Shanghai Jiao Tong University
C
Chuxiao Zeng
Shanghai Jiao Tong University
Bohan Zhao
Bohan Zhao
Scripps Research Institute/HHMI
NeuroscienceMemoryMetabolismSleep
Q
Qinxiang Cao
Shanghai Jiao Tong University