Score
Designs and implements CP-SAT encodings of constraint or combinatorial models and builds verification pipelines that run CP-SAT solvers (e.g., OR-Tools) to determine and certify infeasibility. Produces reproducible computational certificates and automated checks to eliminate remaining obstruction models and verify computer-assisted infeasibility claims.
This work addresses the lack of formal guarantees regarding semantic preservation during problem reformulation and solver correctness in constraint programming. It presents the first end-to-end verified framework implemented in the Lean theorem prover, enabling formal proofs of parameterized equivalence, equisatisfiability, and symmetry-breaking correctness for entire families of problems. The approach combines general, parameterized proofs with instance-level certificate checking, thereby eliminating the need to trust external solvers. Verified certificates are produced via backend transformations, and a single high-level proof suffices for arbitrarily large instances. This methodology achieves dramatic search-space reductions—up to a factor of twenty million—and enables full verification of the largest instances in just a few minutes.
This work addresses the challenge of verifying equivalence between target representations—such as decision-DNNF—and their original CNF encodings in knowledge compilation. Methodologically, it introduces (1) Partitioned-Operation Graphs (POGs) as a unified intermediate representation; (2) the Certified POG (CPOG) proof framework, enabling structured, correctness-preserving compilation from CNF to POG; and (3) full formal verification in Lean 4 of the compiler, proof generator, and model counter. Contributions include: the first end-to-end, machine-checked correctness guarantee for the entire knowledge compilation pipeline; automated verification of D4-generated POGs; empirical evaluation on standard model counting benchmarks; and the first mathematically verified toolchain supporting both weighted and unweighted model counting. The framework ensures semantic equivalence at every compilation step, thereby bridging the gap between practical knowledge compilation tools and formal correctness guarantees.
Correctness verification of MaxSAT solvers has long lacked effective proof mechanisms, particularly within branch-and-bound frameworks where complex inferences—such as look-ahead strategies and pseudo-Boolean constraints encoded via multi-valued decision diagrams (MDDs)—resist generation of checkable certificates. Method: This paper introduces the first systematic extension of proof logging to an advanced branch-and-bound MaxSAT solver, MaxCDCL, supporting full verifiability for clausal encodings, MDD-based constraint representations, and look-ahead reasoning. We propose a low-overhead proof logging mechanism enabling end-to-end certificate generation. Contribution/Results: Our approach bridges the technical gap between MaxSAT’s optimization semantics and formal proof systems, enabling efficient and practical certificate generation. Experimental evaluation confirms its feasibility and scalability, significantly enhancing result trustworthiness. This work establishes the first formal verification foundation for high-assurance combinatorial optimization solvers.
MILP solvers’ outputs lack trustworthy verification in critical applications such as hardware verification, compiler optimization, and machine-assisted theorem proving. Method: This paper proposes the first formal verification framework for VIPR 1.0—a general-purpose certificate format—by fully encoding its inference rule system into unambiguous, SMT-expressible first-order logic formulas and constructing a solver-agnostic verifier compliant with the SMT-LIB standard to ensure algorithmic verifiability. Contribution/Results: The framework eliminates ambiguities inherent in the original VIPR specification and enables rigorous, implementation-independent validation of MILP certificates. Experimental evaluation on public benchmark suites confirms the verifier’s correctness and practical feasibility, demonstrating substantial improvements in both the rigor and generality of MILP certificate verification.
Current benchmarks for mathematical reasoning predominantly rely on answer matching, which fails to assess the logical correctness of solution processes. This work proposes a hybrid verification pipeline that integrates automated and interactive validation by leveraging structured prompting to guide large language models in generating verifiable solutions. The framework supports both formal and informal reasoning and interfaces with proof assistants such as Lean 4, enabling even small-scale models (≤8B parameters) to participate effectively in collaborative verification. Through a multi-agent architecture and advanced prompt engineering, the approach substantially reduces false positive rates. Experimental results demonstrate high verification accuracy across multiple datasets, and the codebase along with deployment guidelines has been publicly released.
This work addresses the challenge of efficiently and reliably importing large-scale logical certificates produced by SAT solvers into Lean 4 to formally verify the unsatisfiability of combinatorial problems. We present the first reflection-based LRAT checker implemented in Lean 4, which directly translates DIMACS formulas and LRAT certificates into Lean theorems without explicitly constructing massive proof terms. Our approach fully supports the internal composition of cube-and-conquer strategies and automatically synthesizes coverage completeness proofs. It significantly outperforms Mathlib’s existing proof-import mechanisms and achieves performance on par with the external checker cake_lpr when verifying large-scale instances such as the Schur number S(4)=44 and the Ramsey number R(4,4)=18, thereby enabling scalable and highly trustworthy automated verification of combinatorial theorems.
This study investigates the feasibility and computational complexity of the edge-cover problem on control flow graphs under semantic constraints. Focusing on five constraint types—POSITIVE, NEGATIVE, ONCE, MAX ONCE, and ALWAYS—it systematically characterizes their complexity landscape through polynomial-time constructions, NP-completeness reductions via SAT variants, and fixed-parameter tractability (FPT) analysis. The results show that the problem is polynomial-time decidable only under POSITIVE constraints; for the other four constraint types, it remains NP-complete even on directed acyclic graphs. Notably, the NEGATIVE-constrained variant admits an FPT algorithm when parameterized by the number of constraints. This work establishes a theoretical foundation and provides algorithmic insights for constrained test generation in program analysis.
This work addresses the lack of efficient and trustworthy verification mechanisms for CTL model checking with fairness constraints by presenting the first self-certifying symbolic model checker that supports interactive certification. The approach leverages Binary Decision Diagrams (BDDs) to perform symbolic verification of CTL properties and introduces, for the first time, an interactive proof system that formally certifies verification results with user-configurable high confidence after solving. By integrating CTL semantics with fairness constraints, QBF solving techniques, and interactive certification, this method preserves full CTL model checking capabilities while delivering a reliable and verifiable automated verification guarantee.
This work presents the first successful application of a large language model—specifically, Claude Opus 4.6—as an AI-powered programming assistant to automatically generate and verify a semantic-preserving proof for the Administrative Normal Form (ANF) transformation in the CertiCoq compiler, entirely without manual proof coding. Guided by human oversight and built upon the Rocq proof language, the approach adapts and transfers techniques from an existing continuation-passing style (CPS) transformation proof. Within approximately 96 hours, the system produced 7,800 lines of machine-checkable proof code, surpassing the previous CPS proof of 5,300 lines and substantially reducing development time. This result demonstrates the feasibility and significant potential of large language models in formal verification.