Learning to Prove, Not Just to Answer: Reinforcement Learning from Formal Verification for Natural-Language Logical Reasoning

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of machine-verifiability in intermediate steps of existing logical reasoning methods, which often leads to erroneous reasoning receiving rewards. To this end, we propose Proof-R1, a framework that integrates reinforcement learning with UNSAT-based formal verification to train large language models for generating verifiable natural language proofs. Specifically, Proof-R1 enforces formal verification at each reasoning step to satisfy proof obligations and restores answer-dependent closures to precisely align reward signals. Experimental results demonstrate that Proof-R1 significantly improves both answer accuracy and process verifiability across three benchmarks and four backbone models, consistently outperforming existing approaches.
📝 Abstract
Large language models (LLMs) are increasingly deployed for natural-language logical reasoning, where the final answer is easy to check but the proof behind it is not. In natural-language logical reasoning, an intermediate conclusion should follow from its premises, and the resulting derivation should support the final answer. Existing methods lack machine-checkable verification of intermediate conclusions and answer-supporting proof dependencies, so they may assign credit to invalid or answer-irrelevant steps. We propose Proof-R1, an RL framework from formal verification that trains LLMs to construct verifiable proofs for natural-language logical reasoning. Proof-R1 admits a generated conclusion into the verified proof state only when the corresponding reasoning action satisfies the proof obligations through UNSAT-based machine-checkable formal verification. Proof-R1 also recovers the answer-supporting dependency closure to trace the proof structure of the final answer and align outcome credit with the proof dependencies. Experiments demonstrate that Proof-R1 improves answer accuracy across three logical reasoning benchmarks and four backbone models and outperforms training-free agents and training-based methods in terms of reasoning-process verifiability.
Problem

Research questions and friction points this paper is trying to address.

Natural-language logical reasoning
Formal verification
Intermediate conclusions
Proof dependencies
Credit assignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Formal Verification
Reinforcement Learning
Logical Reasoning
Proof Dependency
Large Language Models