Score
Design and formally verify translations from problem specifications into conjunctive normal form (CNF), ensuring the encoding is sound and complete so that every CNF solution corresponds to a solution of the original specification and vice versa. Build and analyze both forward encodings and CNF-to-problem mappings and produce machine-checkable proofs or verification artifacts that allow lifting CNF-level results back to the original problem.
This work addresses the problem of verbose and semantically opaque partial assignments induced by CNF conversion in SAT/SMT enumeration. We systematically evaluate the suitability of Tseitin versus Plaisted–Greenbaum (PG) encodings in enumeration contexts. Theoretically and empirically, we show that Tseitin encoding inherently impedes generation of short partial assignments, whereas PG encoding—when combined with negation normal form (NNF) preprocessing—guarantees that each enumerated solution corresponds to a minimal, semantically transparent partial assignment. This synergistic approach is the first to provably eliminate encoding-induced assignment redundancy in enumeration. Evaluated on SMT-LIB benchmarks, it reduces both the number of partial solutions and total runtime by 1–3 orders of magnitude, demonstrating strong theoretical soundness and practical efficacy.
This work presents the first successful application of a large language model—specifically, Claude Opus 4.6—as an AI-powered programming assistant to automatically generate and verify a semantic-preserving proof for the Administrative Normal Form (ANF) transformation in the CertiCoq compiler, entirely without manual proof coding. Guided by human oversight and built upon the Rocq proof language, the approach adapts and transfers techniques from an existing continuation-passing style (CPS) transformation proof. Within approximately 96 hours, the system produced 7,800 lines of machine-checkable proof code, surpassing the previous CPS proof of 5,300 lines and substantially reducing development time. This result demonstrates the feasibility and significant potential of large language models in formal verification.
This work addresses the challenge of verifying equivalence between target representations—such as decision-DNNF—and their original CNF encodings in knowledge compilation. Methodologically, it introduces (1) Partitioned-Operation Graphs (POGs) as a unified intermediate representation; (2) the Certified POG (CPOG) proof framework, enabling structured, correctness-preserving compilation from CNF to POG; and (3) full formal verification in Lean 4 of the compiler, proof generator, and model counter. Contributions include: the first end-to-end, machine-checked correctness guarantee for the entire knowledge compilation pipeline; automated verification of D4-generated POGs; empirical evaluation on standard model counting benchmarks; and the first mathematically verified toolchain supporting both weighted and unweighted model counting. The framework ensures semantic equivalence at every compilation step, thereby bridging the gap between practical knowledge compilation tools and formal correctness guarantees.
Automatically verifying C programs generated by large language models (LLMs) remains challenging due to their syntactic and semantic irregularities, which hinder formal verification. Method: This paper proposes SynVer—a novel framework that tightly integrates LLM-based program synthesis with formal verification. SynVer introduces verifiability-aware biasing mechanisms operating at both syntactic and semantic levels to guide LLMs toward generating verification-friendly code. It further incorporates separation logic (SL) specifications and the Verified Software Toolchain (VST) to enable end-to-end, fully automated verification—from specification to C implementation to machine-checked safety proofs. Results: Evaluated on diverse benchmarks covering basic coding tasks, SL assertions, and API specifications, SynVer significantly improves the automatic verification success rate of LLM-generated C programs. Empirical results demonstrate its scalability, robustness, and effectiveness in bridging the gap between neural code generation and rigorous formal assurance.
This work addresses the absence of solver- and domain-agnostic verification mechanisms in Semantic-Guided Synthesis (SemGuS). Methodologically, it rigorously reduces correctness checking of SemGuS solutions to validity checking in Constraint Logic Programming (CLP), uniformly supporting first-order logic, constrained/coinstrained Horn clauses, and general CLP queries; it further extends the SemGuS syntax to accommodate nondeterministic and reactive synthesis. The key contributions are: (i) the first sound and complete reduction of SemGuS verification to CLP validity, thereby overcoming prior expressiveness limitations; (ii) enabling verification of previously inexpressible complex synthesis instances; and (iii) integration into an enumerative solver, achieving successful synthesis on benchmark problems unsolved by all existing SemGuS solvers. This framework establishes a foundational, general-purpose verification infrastructure for SemGuS.
This work presents the first systematic solution to the problems of model counting and sampling in the theory of bit-vectors. By leveraging bit-blasting to translate bit-vector formulas into conjunctive normal form (CNF), the approach integrates modern CNF counters and samplers to uniformly support a wide range of modes, including exact and approximate counting, projected and unprojected counting, as well as near-uniform and uniform sampling. The resulting tool, csb, addresses a critical gap in efficient counting and sampling for bit-vector constraints and demonstrates substantial performance advantages over existing methods in empirical evaluations, highlighting its practicality and effectiveness.
This work addresses the lack of formal guarantees regarding semantic preservation during problem reformulation and solver correctness in constraint programming. It presents the first end-to-end verified framework implemented in the Lean theorem prover, enabling formal proofs of parameterized equivalence, equisatisfiability, and symmetry-breaking correctness for entire families of problems. The approach combines general, parameterized proofs with instance-level certificate checking, thereby eliminating the need to trust external solvers. Verified certificates are produced via backend transformations, and a single high-level proof suffices for arbitrarily large instances. This methodology achieves dramatic search-space reductions—up to a factor of twenty million—and enables full verification of the largest instances in just a few minutes.
This work addresses the inefficiency of query operations—such as uniform sampling, direct access, and model enumeration—on propositional formulas in conjunctive normal form (CNF) after compilation into deterministic decomposable negation normal form (d-DNNF). To overcome this limitation, the authors propose a preprocessing technique that preserves only the model count rather than full logical equivalence. Applied prior to CNF-to-d-DNNF compilation, this method optimizes the input formula while retaining essential preprocessing information to accelerate downstream queries. The study presents the first systematic evaluation of model-count-preserving preprocessors, demonstrating their ability to substantially enhance the performance of diverse query tasks on d-DNNF representations, thereby surpassing the constraints of traditional equivalence-preserving preprocessing. Extensive experiments across multiple benchmark domains confirm the approach’s efficiency and robustness.
This work addresses the state-space explosion problem in formal verification of large-scale C programs by proposing a novel approach that integrates large language models (LLMs) with compositional verification. The method leverages an LLM to automatically generate function contracts from system-level specifications and coordinates system-wide and function-level verification within a CEGAR-CEGIS loop. A SMART ICE learning mechanism refines these contracts whenever verification fails. This is the first framework to incorporate LLMs into hierarchical contract-based verification, substantially reducing the number of refinement iterations and enhancing scalability. Experimental results demonstrate success rates of 82–96% on Frama-C, 33–50% on X.509, 82–88% on LF2C-Simple, 55–64% on VerifyThis, and 67% on LF-Hard benchmarks, with 93–95% of Frama-C programs verified in just a single iteration.