🤖 AI Summary
This study addresses the limitation of existing formal reasoning benchmarks, which evaluate only complete proofs and lack fine-grained, step-level diagnostic capabilities. We propose a novel method that generates localized proof obligations from natural language solutions. By leveraging a natural language processing pipeline and semantic alignment techniques, mathematical analysis exercises are decomposed into context- and goal-isolated proof gaps. Based on the Demidovich problem set, we construct a fine-grained benchmark dataset comprising 3,015 exercises and 26,116 proof gaps, achieving a breakthrough in evaluation granularity from the theorem level to the step level. This benchmark precisely localizes model failure points during formal proof construction, significantly enhancing diagnostic precision and providing foundational support for the future development of proof verification systems.
📝 Abstract
Existing formal mathematics benchmarks, such as miniF2F, ProofNet, and PutnamBench, primarily evaluate models on constructing complete formal proofs for challenging problems. Because success is measured at the theorem level, these benchmarks offer limited insight into models' step-level formal reasoning. Evaluating this capability separately enables finer-grained diagnosis of model limitations than theorem-level evaluation alone.
To fill this evaluation gap, we introduce ProofGap, a fine-grained benchmark for step-level formal reasoning. ProofGap is constructed through a natural-language proof-processing pipeline that decomposes each reasoning step into one or more aligned proof gaps. Applying this pipeline to natural-language solutions to 3,015 exercises in B. P. Demidovich's Problems in Mathematical Analysis yields 26,116 gaps. The benchmark focuses on mathematical analysis, a domain that remains challenging for current models. By supplying the local context and target explicitly, gap completion isolates local formal proof construction from end-to-end proof composition, enabling more precise localization of model failures. Natural-language solutions serve as the provenance of these obligations, while the benchmark task itself starts from an already formalized local context and goal. Beyond benchmarking, the same pipeline may support future proof-verification systems, provided that semantic translation and sequential proof composition are handled reliably.