Score
Designs and implements automated checkers and end-to-end verification pipelines that validate formal certificates and proofs, including mechanical certificate checkers and Python verification pipelines that recompute and certify claimed parameters. Builds verifiers and validators that run polynomial-time checks, confirm constraint satisfaction, frame NP-membership arguments, and produce reproducible, auditable certificate validation outputs (e.g., for reductions or hardness claims).
This work addresses the challenge of reliably importing VeriPB proof certificates generated by pseudo-Boolean (PB) solvers into Lean 4 and achieving end-to-end trustworthy verification from solver outputs back to the semantics of the original combinatorial problem. To this end, we present the first formalization in Lean 4 of a reflective proof checker that fully supports the VeriPB kernel rules. By integrating native code compilation, verified encoding transformations, and cutting-plane derivation techniques, our approach efficiently handles large-scale proofs comprising tens of thousands of inference steps while avoiding the memory bottlenecks associated with explicit proof-term construction. This method bridges the trust gap between solver output and formal semantics, producing reusable and composable Lean theorems, and demonstrates both effectiveness and scalability across multiple combinatorial problems.
Formal verification remains inaccessible to most researchers due to the high barrier of expertise and the gap between theoretical algorithms described in academic papers and executable, verifiable code. Method: This paper introduces the first end-to-end framework that automatically extracts proof structures from formally verified research papers, models them as logical specifications, and generates provably correct code compatible with theorem provers (e.g., Coq, Isabelle). The approach integrates Transformer-based fine-grained text understanding, structured reasoning, and formal-verification–aware prompt engineering to reconstruct omitted low-level proof details. Contribution/Results: Evaluated on canonical domains—including distributed consensus and cryptographic protocols—the generated code achieves both compilability and foundational verifiability. By bridging the gap between peer-reviewed formal proofs and machine-checkable implementations, this work substantially lowers the entry barrier to formal verification and establishes a novel paradigm for constructing high-assurance systems directly from academic literature.
This work addresses the challenge of achieving efficient and reliable formal verification of production-grade cryptographic code written in Rust. We present the first end-to-end Rust-to-Lean 4 verification pipeline, integrating the Charon, Aeneas, and Hax frameworks for symbolic extraction, leveraging the ArkLib and CompPoly libraries of formally specified cryptographic primitives, and introducing the Aristotle and Aleph AI-powered provers to automatically discharge complex proof obligations. All results are rigorously validated by the Lean 4 kernel. Our approach successfully reproduces and fully verifies key cryptographic primitives from Plonky3 and RISC Zero—including FRI folding, finite field arithmetic, Horner evaluation, and Merkle inclusion proofs—and automatically completes proofs for two longstanding open conjectures.
This work addresses the critical challenge of reliably integrating automated reasoning tools—such as theorem provers, SAT/SMT solvers, and termination analyzers—with proof assistants to build highly trustworthy systems. It presents a systematic survey and comparative analysis of two principal technical approaches: certification and formal verification. The study examines core methodologies including logical encoding, result replay and checking, and integration mechanisms within proof assistants. By elucidating the respective strengths and limitations of these methods and illustrating them through multiple successful case studies, the paper offers clear methodological guidance for constructing high-assurance automated reasoning systems, thereby substantially enhancing the verifiability and trustworthiness of their outputs.
This work addresses the long-standing absence of explicit deterministic polynomial-time solvers for concrete NP-complete problems by presenting the first complete deterministic Turing machine implementation for specific problems such as SAT and Subset-Sum. Building upon an NP-verifier simulation framework, the approach extends verification mechanisms to deterministic FNP solving—without increasing the polynomial time complexity—through techniques including dynamic computation graphs, feasible graph construction, and verification-path traversal. A fully functional simulator is implemented in Python, and experimental results demonstrate that the system strictly adheres to polynomial time bounds while effectively generating valid witnesses for satisfiable instances. The source code is publicly released to ensure transparency and reproducibility of the results.
This work addresses the challenge of efficiently verifying whether an AI agent’s behavior adheres to a prescribed policy without relying on trust in the agent or re-executing its computations. The authors propose a novel paradigm that integrates formal methods with cryptographic proofs: policy specifications are encoded as logical predicates, compiled into polynomial constraints, and used to generate succinct, independently verifiable certificates via Succinct Non-interactive Arguments of Knowledge (SNARKs), optionally with zero-knowledge guarantees. This framework enables an end-to-end transformation from high-level policy statements to verifiable evidence, facilitating trustless compliance auditing and bridging the gap between AI governance, deployment, and formal verification.
This work addresses the challenge of ensuring determinism and trustworthiness in structured computations surrounding large language models (LLMs) without directly verifying the LLMs themselves. It proposes a trust-boundary architecture grounded in Lean 4, leveraging formally verifiable certificates to rigorously certify structured components within LLM pipelines. The core innovation comprises three families of local certificates and two composition operators, enabling hierarchical assumption management, extraction of maximal certifiable residuals, and closed-form computation of end-to-end perturbation budgets. The approach integrates Lean 4 kernel type checking, axiom-free auditing (with 17 out of 46 assertions proven without axioms), dual-lattice grounding, sensitivity analysis, and Hoare-style action logic, formally covering 22 certificate types. Empirical validation is demonstrated across four scenarios: HotpotQA reasoning, embedding stability, file system agents, and related structured tasks.
This work addresses the lack of formal guarantees regarding semantic preservation during problem reformulation and solver correctness in constraint programming. It presents the first end-to-end verified framework implemented in the Lean theorem prover, enabling formal proofs of parameterized equivalence, equisatisfiability, and symmetry-breaking correctness for entire families of problems. The approach combines general, parameterized proofs with instance-level certificate checking, thereby eliminating the need to trust external solvers. Verified certificates are produced via backend transformations, and a single high-level proof suffices for arbitrarily large instances. This methodology achieves dramatic search-space reductions—up to a factor of twenty million—and enables full verification of the largest instances in just a few minutes.
This work addresses subtle vulnerabilities in zkEVM implementations—such as incorrect gas computations—that can yield semantically flawed yet formally valid zero-knowledge proofs, thereby compromising system security. Existing formal verification approaches rely on manually crafted specifications, limiting their scalability. To overcome this, the paper introduces VeriSynth, a novel hybrid framework that uniquely combines large language models (LLMs) as a formalization frontend with SMT solvers as correctness oracles. Through semantic decomposition, retrieval-augmented prompting, and verification-guided self-repair, VeriSynth automatically translates Rust-based zkEVM opcodes into executable symbolic constraint models in Python/Z3, enabling closed-loop constraint synthesis and repair. Evaluated on the first source-level zkEVM verification benchmark, VeriSynth achieves over 90% vulnerability detection accuracy, substantially outperforming pure LLM approaches, conversational baselines, and industrial-grade handcrafted test suites, with ablation studies confirming the necessity of each component.
This work addresses the challenges of scale and complexity in formally verifying production-grade cryptographic libraries, where existing approaches fall short of end-to-end automation. We present CryptoProver, a system that achieves, for the first time, fully automated verification of real-world cryptographic implementations such as curve25519-dalek and RustCrypto’s chacha20. CryptoProver integrates large language models with the Verus verifier to automatically synthesize internal specifications and verifiable proofs from high-level API contracts, without requiring source code modifications. By leveraging a pre-defined trusted library, mechanical gating, and isolation mechanisms, the system ensures specification strength and cross-module consistency. In experiments, CryptoProver completed verification within 11.4 hours at an API cost of \$466.99, successfully covering core cryptographic components relied upon by widely deployed systems including Signal (with 218 million downloads) and Shadowsocks.