ProofLoom: Proof-Obligation-Driven Theory Construction for Autoformalizing Research-Level Stochastic Optimization

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the theoretical drift caused by model modifications and the challenges of automated construction in the formalization of stochastic optimization algorithms. We propose a fully automated formalization framework driven by large language model (LLM) agents. This framework employs proof obligations to guide the automatic construction of Lean models and supporting theories, introduces signature contracts alongside independent auditing mechanisms to prevent assumption weakening, and establishes a reusable verification library, SOptLib, to enable cumulative verification cycles. Experimental results demonstrate that the system achieves an average score of 6.3 out of 7 across 15 tasks, generates 490,000 lines of `sorry`-free code, and identifies 28 formula errors and proof gaps in published literature, thereby realizing highly reliable automated formalization of research-grade algorithms.
📝 Abstract
Formalizing research-level stochastic optimization in Lean requires both an algorithm model and domain theory connecting foundational libraries to convergence proofs. Revising a model to restore provability can change the mathematical claim. We introduce ProofLoom, a fully automated LLM-agent system for Proof-Obligation-Driven Theory Construction. Given a published algorithm, target theorem, and source proof, ProofLoom autonomously constructs the Lean model and supporting theory. Open proof obligations drive the development of definitions, interfaces, lemmas, and proof plans. Signature contracts record evidence and obligations for model revisions; an independent Judge rejects unsupported assumptions and weakened conclusions. Planner expands the published argument into intermediate claims, and Audit checks whether the Lean proof follows it. Across tasks, SOptLib accumulates verified mathematics and construction experience: reusable results are extracted, generalized, and verified, while modeling decisions and failed proof routes are recorded. Later tasks retrieve these results and records and contribute new developments, forming a cycle of construction, accumulation, and reuse. On fifteen textbook and research-paper tasks, ProofLoom obtains mean human ratings of 6.3/7 and 6.4/7, compared with 4.9/7 and 5.0/7 for the strongest of six baselines. Across 33 developments, it produces 490,693 lines of algorithm-local Lean code with no sorry. The formalizations also expose 28 incorrect formulas, proof gaps, and algorithm-analysis mismatches in published sources across 22 developments, each with checked evidence.
Problem

Research questions and friction points this paper is trying to address.

stochastic optimization
autoformalization
proof obligations
Lean
domain theory
Innovation

Methods, ideas, or system contributions that make the work stand out.

Autoformalization
Proof-Obligation-Driven
LLM-Agent
Stochastic Optimization
Theory Construction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.