Score
Designs and implements systems that check and certify candidate solutions against hard constraints, using constraint‑programming/CP‑SAT solvers (e.g., Google OR‑Tools CP‑SAT) to verify satisfiability and produce validation artifacts; when candidates are infeasible, builds automated repair procedures or rejection filters that modify or discard candidates so they meet required feasibility conditions (such as hard margins or contractual terms) before recommendations are reported.
This work addresses the lack of formal guarantees regarding semantic preservation during problem reformulation and solver correctness in constraint programming. It presents the first end-to-end verified framework implemented in the Lean theorem prover, enabling formal proofs of parameterized equivalence, equisatisfiability, and symmetry-breaking correctness for entire families of problems. The approach combines general, parameterized proofs with instance-level certificate checking, thereby eliminating the need to trust external solvers. Verified certificates are produced via backend transformations, and a single high-level proof suffices for arbitrarily large instances. This methodology achieves dramatic search-space reductions—up to a factor of twenty million—and enables full verification of the largest instances in just a few minutes.
Manual construction of reductions among NP-complete problems is error-prone and difficult to verify. To address this, this paper proposes a formal, SAT-based methodology for modeling and automatically verifying reductions—marking the first systematic application of SAT solving (via the URSA constraint-solving system) to encoding reduction specifications, logical analysis, and correctness proofs. The approach translates reduction constructions into Boolean constraint models, enabling automated reasoning and counterexample detection. Experimental evaluation on canonical NP-complete reductions—including 3SAT → Vertex Cover and 3SAT → Hamiltonian Cycle—demonstrates that the method efficiently generates machine-checkable correctness proofs and successfully uncovers several previously undetected subtle errors in the literature. Consequently, it significantly enhances the accuracy, reliability, and scalability of reduction verification.
This study addresses the high cost and error-proneness of model refactoring caused by paradigm disparities among constraint solvers. We propose a modular automated translation framework based on CPMpy that employs a layered waterfall architecture to uniformly handle sub-expression negation and auxiliary variable generation while optimizing linearization strategies for nonlinear operators. This approach enables seamless translation from high-level models to low-level paradigms, including CP, SMT, and ILP. Experimental results demonstrate that the framework effectively eliminates manual rewriting and that its optimized linearization significantly enhances ILP and PB solving performance. Consequently, this work provides an efficient, flexible, and standardized solution for the automatic benchmarking of multi-paradigm solvers in combinatorial optimization.
Correctness verification of MaxSAT solvers has long lacked effective proof mechanisms, particularly within branch-and-bound frameworks where complex inferences—such as look-ahead strategies and pseudo-Boolean constraints encoded via multi-valued decision diagrams (MDDs)—resist generation of checkable certificates. Method: This paper introduces the first systematic extension of proof logging to an advanced branch-and-bound MaxSAT solver, MaxCDCL, supporting full verifiability for clausal encodings, MDD-based constraint representations, and look-ahead reasoning. We propose a low-overhead proof logging mechanism enabling end-to-end certificate generation. Contribution/Results: Our approach bridges the technical gap between MaxSAT’s optimization semantics and formal proof systems, enabling efficient and practical certificate generation. Experimental evaluation confirms its feasibility and scalability, significantly enhancing result trustworthiness. This work establishes the first formal verification foundation for high-assurance combinatorial optimization solvers.
MILP solvers’ outputs lack trustworthy verification in critical applications such as hardware verification, compiler optimization, and machine-assisted theorem proving. Method: This paper proposes the first formal verification framework for VIPR 1.0—a general-purpose certificate format—by fully encoding its inference rule system into unambiguous, SMT-expressible first-order logic formulas and constructing a solver-agnostic verifier compliant with the SMT-LIB standard to ensure algorithmic verifiability. Contribution/Results: The framework eliminates ambiguities inherent in the original VIPR specification and enables rigorous, implementation-independent validation of MILP certificates. Experimental evaluation on public benchmark suites confirms the verifier’s correctness and practical feasibility, demonstrating substantial improvements in both the rigor and generality of MILP certificate verification.
This work addresses the critical challenge of reliably integrating automated reasoning tools—such as theorem provers, SAT/SMT solvers, and termination analyzers—with proof assistants to build highly trustworthy systems. It presents a systematic survey and comparative analysis of two principal technical approaches: certification and formal verification. The study examines core methodologies including logical encoding, result replay and checking, and integration mechanisms within proof assistants. By elucidating the respective strengths and limitations of these methods and illustrating them through multiple successful case studies, the paper offers clear methodological guidance for constructing high-assurance automated reasoning systems, thereby substantially enhancing the verifiability and trustworthiness of their outputs.
This study addresses the difficulty coding agents face in verifying whether programs satisfy specifications and their inability to effectively leverage expert diagnostic experience from verification failures. Building upon the executable semantics of the K framework, this work proposes a method that transforms expert diagnostics into reusable guidance, establishing a comprehensive verification pipeline encompassing specification generation, proof repair, and adequacy auditing. By designing paired clean and defective program packages, the effectiveness of the auditing mechanism is systematically evaluated. The proposed approach achieves a 164/164 pass rate on the HumanEval benchmark, with all defects accurately identified. Furthermore, experiments on KleverBench reveal optimization opportunities for guidance selection strategies under resource constraints. This research offers a novel paradigm for enhancing the formal verification capabilities of coding agents.
This work addresses the challenge of efficiently and reliably importing large-scale logical certificates produced by SAT solvers into Lean 4 to formally verify the unsatisfiability of combinatorial problems. We present the first reflection-based LRAT checker implemented in Lean 4, which directly translates DIMACS formulas and LRAT certificates into Lean theorems without explicitly constructing massive proof terms. Our approach fully supports the internal composition of cube-and-conquer strategies and automatically synthesizes coverage completeness proofs. It significantly outperforms Mathlib’s existing proof-import mechanisms and achieves performance on par with the external checker cake_lpr when verifying large-scale instances such as the Schur number S(4)=44 and the Ramsey number R(4,4)=18, thereby enabling scalable and highly trustworthy automated verification of combinatorial theorems.
This study addresses the longstanding challenge of lacking formal connections between search and refutation algorithms for random constraint satisfaction problems (CSPs). By leveraging average-case complexity theory and semi-random model analysis, this work constructs a unified theoretical framework that rigorously relates search and refutation processes for the first time. It demonstrates that existing algorithms can output near-optimal solutions accompanied by certificates, and proposes optimization methods with verifiable ε-optimality guarantees. Furthermore, novel robust algorithms are designed to operate under strong contamination models. Ultimately, this research achieves verifiably near-optimal solving for semi-random and contaminated CSPs, effectively unifying computational threshold theory and providing a new paradigm that combines theoretical rigor with practical utility for related fields.
This study addresses the absence of a unified inference calculus for Constrained Horn Clause (CHC) verification and its unclear relationship with symbolic execution by proposing a constrained resolution calculus framework. This approach formalizes both forward and backward symbolic execution as special cases of the calculus, enabling complete reasoning for linear and nonlinear CHCs. By establishing the refutational completeness of the resolution strategy, the method generalizes k-induction to nonlinear CHCs and introduces criteria for pruning redundant derivations. Experimental evaluations based on the Eldarica solver demonstrate that the proposed framework exhibits verification capabilities complementary to existing CEGAR- and IC3-based techniques.