Advocate for Complete Benchmarks for Formal Reasoning with Formal/Informal Statements and Formal/Informal Proofs

πŸ“… 2025-07-07
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF

career value

195K/year
πŸ€– AI Summary
Current benchmarks in formal reasoning and automated theorem proving suffer from incomplete coverage, erroneous annotations, and closed-source code and dataβ€”leading to unreliable evaluations, poor reproducibility, and hindered community collaboration. To address these issues, we propose an end-to-end open-source benchmark construction framework: (1) unifying formal and informal statements and proofs; (2) rigorously verifying logical correctness and domain coverage; (3) fully open-sourcing benchmark datasets, evaluation scripts, and baseline models; and (4) systematically identifying and rectifying misleading evaluation practices (e.g., data leakage, undetected overfitting). Our core contribution is the first standardized evaluation suite that simultaneously ensures completeness, verifiability, and openness. This significantly improves result comparability and reproducibility, lowers barriers to entry, and enables fair, cross-method, and cross-community benchmarking and collaborative innovation.

Technology Category

Application Category

πŸ“ Abstract
This position paper provides a critical but constructive discussion of current practices in benchmarking and evaluative practices in the field of formal reasoning and automated theorem proving. We take the position that open code, open data, and benchmarks that are complete and error-free will accelerate progress in this field. We identify practices that create barriers to contributing to this field and suggest ways to remove them. We also discuss some of the practices that might produce misleading evaluative information. We aim to create discussions that bring together people from various groups contributing to automated theorem proving, autoformalization, and informal reasoning.
Problem

Research questions and friction points this paper is trying to address.

Advocate for complete error-free benchmarks in formal reasoning
Identify barriers to contribution in automated theorem proving
Address misleading evaluative practices in formal reasoning benchmarks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Promote complete error-free benchmarks
Advocate open code and data
Encourage collaborative discussions
πŸ”Ž Similar Papers
No similar papers found.