Score
Design and implement verifiers and parsers that import CNF formulas and LRAT (Linear RAT) proof certificates and validate each RAT-style inference step to confirm propositional refutations. Build standalone checkers or integrations into proof assistants that replay LRAT refutations, enforce correctness of linear RAT operations, and emit machine-checkable verification artifacts without relying on external SAT solvers.
This work addresses the challenge of efficiently and reliably importing large-scale logical certificates produced by SAT solvers into Lean 4 to formally verify the unsatisfiability of combinatorial problems. We present the first reflection-based LRAT checker implemented in Lean 4, which directly translates DIMACS formulas and LRAT certificates into Lean theorems without explicitly constructing massive proof terms. Our approach fully supports the internal composition of cube-and-conquer strategies and automatically synthesizes coverage completeness proofs. It significantly outperforms Mathlib’s existing proof-import mechanisms and achieves performance on par with the external checker cake_lpr when verifying large-scale instances such as the Schur number S(4)=44 and the Ramsey number R(4,4)=18, thereby enabling scalable and highly trustworthy automated verification of combinatorial theorems.
This work addresses key limitations of large language models in automated formal verification—namely, scarce training data, difficulty adhering to strictly checkable specifications, and a tendency to exploit weak specifications through heuristic shortcuts. To mitigate these issues, the authors propose a recursive reasoning framework that integrates Group Relative Policy Optimization (GRPO) with verifier-guided feedback. The approach features multi-round reinforcement learning, verifier-driven reasoning scaffolds, and mechanisms for subgoal decomposition and proof revision. This methodology substantially reduces specification misuse and significantly improves correctness rates: on a refined Dafny benchmark, verification success rises from 9.7% to 31.1%; in Lean, combining proof revision boosts VeriCoding pass rates from 46.2% to 69.2%, and seven previously unsolved tasks in VERINA are successfully resolved.
In program verification, SMT solvers frequently fail due to missing critical assertions, necessitating manual assertion hints and substantially increasing verification overhead. This paper proposes an LLM-based automated assertion completion method: (1) it precisely localizes assertion gaps using SMT error messages and introduces placeholder tokens; (2) it defines a code-level proof similarity metric to enable context-aware example retrieval; and (3) it integrates domain-specific prompt engineering with SMT feedback-driven iterative refinement. Evaluated on the DafnyGym benchmark, our approach generates over 56.6% of required assertions in a single attempt, significantly improving automated verification success rates. Our key contributions are the first integration of error-driven localization, proof-aware retrieval, and verification-feedback closed-loop optimization into an LLM-assisted assertion generation framework.
Existing evaluation methods treat Rust program verification as a black box, relying solely on binary outcomes to assess whether large language models (LLMs) generate valid proof hints, thereby failing to capture their logical reasoning capabilities. This work proposes VCoT-Lift, a novel framework that, for the first time, lifts the low-level reasoning traces of automated theorem provers into human-readable, high-level “verification chains of thought.” To enable fine-grained assessment, we construct VCoT-Bench—a benchmark comprising 1,988 tasks—evaluating LLMs along three dimensions: robustness to missing proofs, type coverage, and positional sensitivity. Our evaluation of ten state-of-the-art LLMs reveals their fragile performance in formal Rust verification, demonstrating a significant gap between current models and the capabilities of dedicated theorem provers.
MILP solvers’ outputs lack trustworthy verification in critical applications such as hardware verification, compiler optimization, and machine-assisted theorem proving. Method: This paper proposes the first formal verification framework for VIPR 1.0—a general-purpose certificate format—by fully encoding its inference rule system into unambiguous, SMT-expressible first-order logic formulas and constructing a solver-agnostic verifier compliant with the SMT-LIB standard to ensure algorithmic verifiability. Contribution/Results: The framework eliminates ambiguities inherent in the original VIPR specification and enables rigorous, implementation-independent validation of MILP certificates. Experimental evaluation on public benchmark suites confirms the verifier’s correctness and practical feasibility, demonstrating substantial improvements in both the rigor and generality of MILP certificate verification.
Traditional formal verification relies on expert-crafted proofs, which are difficult to scale, while existing large language model–based approaches suffer from low coverage due to reliance on predefined proof strategies. This work proposes a novel paradigm that delegates the generation of complete lemma proofs to general-purpose code agents (e.g., Claude Code), guided by a verification framework that enforces hard constraints and provides feedback to guarantee correctness, completeness, and termination. By abandoning fixed proof strategies, this approach achieves, for the first time, fully automated, end-to-end formal verification without human intervention and with full coverage across multiple proof assistants, including Coq and Lean. Experiments demonstrate its effectiveness: it automatically verifies all 4,474 lemmas in Iris logic and the Rust standard library, attains 100% coverage on the reglang benchmark, and successfully proves 72 previously unverified lemmas in iris-lean.
This work addresses the critical challenge of reliably integrating automated reasoning tools—such as theorem provers, SAT/SMT solvers, and termination analyzers—with proof assistants to build highly trustworthy systems. It presents a systematic survey and comparative analysis of two principal technical approaches: certification and formal verification. The study examines core methodologies including logical encoding, result replay and checking, and integration mechanisms within proof assistants. By elucidating the respective strengths and limitations of these methods and illustrating them through multiple successful case studies, the paper offers clear methodological guidance for constructing high-assurance automated reasoning systems, thereby substantially enhancing the verifiability and trustworthiness of their outputs.