🤖 AI Summary
This study addresses the lack of machine-verifiability in intermediate steps of existing logical reasoning methods, which often leads to erroneous reasoning receiving rewards. To this end, we propose Proof-R1, a framework that integrates reinforcement learning with UNSAT-based formal verification to train large language models for generating verifiable natural language proofs. Specifically, Proof-R1 enforces formal verification at each reasoning step to satisfy proof obligations and restores answer-dependent closures to precisely align reward signals. Experimental results demonstrate that Proof-R1 significantly improves both answer accuracy and process verifiability across three benchmarks and four backbone models, consistently outperforming existing approaches.
📝 Abstract
Large language models (LLMs) are increasingly deployed for natural-language logical reasoning, where the final answer is easy to check but the proof behind it is not. In natural-language logical reasoning, an intermediate conclusion should follow from its premises, and the resulting derivation should support the final answer. Existing methods lack machine-checkable verification of intermediate conclusions and answer-supporting proof dependencies, so they may assign credit to invalid or answer-irrelevant steps. We propose Proof-R1, an RL framework from formal verification that trains LLMs to construct verifiable proofs for natural-language logical reasoning. Proof-R1 admits a generated conclusion into the verified proof state only when the corresponding reasoning action satisfies the proof obligations through UNSAT-based machine-checkable formal verification. Proof-R1 also recovers the answer-supporting dependency closure to trace the proof structure of the final answer and align outcome credit with the proof dependencies. Experiments demonstrate that Proof-R1 improves answer accuracy across three logical reasoning benchmarks and four backbone models and outperforms training-free agents and training-based methods in terms of reasoning-process verifiability.