Score
Design and construct verifiable certificates—explicit witness structures or combinatorial objects—that demonstrate and allow efficient verification of a claimed solution by encoding and enforcing required constraints (e.g., triple and LCA constraints) and detecting or handling forbidden configurations.
This work addresses the unreliability of language models in physical design by introducing the Physics-Anchored Certification (PHACT) framework, which shifts certification authority from the model to a deterministic verification engine grounded in physical laws. In this paradigm, the language model serves solely as a proposal mechanism for candidate designs, while formal certification is performed by the engine through structured metrics derived from fixed inputs, thereby eliminating the possibility of fabrication at its source. The resulting propose-and-certify closed-loop system demonstrates robustness across five scientific domains and achieves zero erroneous certifications in 80 adversarial trials—spanning two language models, two decoding temperatures, and deliberately compromised verification engines—significantly enhancing the trustworthiness of generated outputs.
This work addresses the challenge of verifying that confidential software adheres to publicly specified temporal functional properties without revealing its internal implementation. It introduces, for the first time, zero-knowledge proofs into deductive model checking and proposes a novel method capable of generating verifiable correctness certificates. The approach supports both explicit-state transition graphs and symbolic linear guarded command representations of systems, integrating key techniques including polynomial commitments, Farkas’ lemma, piecewise-linear ranking functions, and Sigma protocols—encompassing matrix multiplication and range proofs. A prototype implementation demonstrates the practicality of the method on LTL verification benchmarks, achieving strong formal guarantees while preserving system confidentiality.
This work addresses the lack of formal guarantees regarding semantic preservation during problem reformulation and solver correctness in constraint programming. It presents the first end-to-end verified framework implemented in the Lean theorem prover, enabling formal proofs of parameterized equivalence, equisatisfiability, and symmetry-breaking correctness for entire families of problems. The approach combines general, parameterized proofs with instance-level certificate checking, thereby eliminating the need to trust external solvers. Verified certificates are produced via backend transformations, and a single high-level proof suffices for arbitrarily large instances. This methodology achieves dramatic search-space reductions—up to a factor of twenty million—and enables full verification of the largest instances in just a few minutes.
This work addresses the challenge of ensuring determinism and trustworthiness in structured computations surrounding large language models (LLMs) without directly verifying the LLMs themselves. It proposes a trust-boundary architecture grounded in Lean 4, leveraging formally verifiable certificates to rigorously certify structured components within LLM pipelines. The core innovation comprises three families of local certificates and two composition operators, enabling hierarchical assumption management, extraction of maximal certifiable residuals, and closed-form computation of end-to-end perturbation budgets. The approach integrates Lean 4 kernel type checking, axiom-free auditing (with 17 out of 46 assertions proven without axioms), dual-lattice grounding, sensitivity analysis, and Hoare-style action logic, formally covering 22 certificate types. Empirical validation is demonstrated across four scenarios: HotpotQA reasoning, embedding stability, file system agents, and related structured tasks.
This work addresses the challenge of reliably importing VeriPB proof certificates generated by pseudo-Boolean (PB) solvers into Lean 4 and achieving end-to-end trustworthy verification from solver outputs back to the semantics of the original combinatorial problem. To this end, we present the first formalization in Lean 4 of a reflective proof checker that fully supports the VeriPB kernel rules. By integrating native code compilation, verified encoding transformations, and cutting-plane derivation techniques, our approach efficiently handles large-scale proofs comprising tens of thousands of inference steps while avoiding the memory bottlenecks associated with explicit proof-term construction. This method bridges the trust gap between solver output and formal semantics, producing reusable and composable Lean theorems, and demonstrates both effectiveness and scalability across multiple combinatorial problems.
This work demonstrates a fundamental limitation of unified formal verification methods within the standard Turing model when applied to nontrivial semantic invariants. By formalizing “acceptable verification schemes” as generator–verifier pairs through a model-theoretic lens, the study integrates Rice’s theorem with formal verification frameworks to prove that such schemes implicitly induce undecidable decision procedures. Crucially, this impossibility stems from the computational behavior inherent to the verification mechanism itself, rather than from unprovable complexity-theoretic assumptions. Leveraging computability theory, model theory, and Coq-based formalization, the authors construct an extended structural model capturing semantic–syntactic interactions and rigorously establish that properties related to P vs NP and cryptographic assumptions such as one-way functions cannot be certified by any such unified method. A complete Coq implementation accompanies the theoretical results.
This work addresses a central challenge in system security: formally verifying that system designs and implementations satisfy intended safety properties and support security certification. The authors propose a systematic approach grounded in proof assistants, integrating interactive theorem proving and formal methods to precisely model and machine-check critical security properties across diverse domains—including system security, language-level security, secure compilation, and cryptography. By enabling rigorous, machine-verifiable proofs of correctness, this methodology significantly strengthens the formal assurance of security properties and provides a unified theoretical framework and toolchain for constructing verifiable and certifiable secure systems.
This work addresses the significant disparity in verifiability among semantically equivalent yet structurally diverse programs, a key bottleneck in generating high-assurance software. The authors propose Diversify2Verify, a novel approach that leverages large language models to synthesize diverse recursive and imperative implementations of the same task, integrates the Why3 platform for automatic contract inference and formal verification, and introduces a verifier-guided annotation repair mechanism to enhance verifiability. This study is the first to systematically expose the verifiability gap across equivalent program variants and establishes a new paradigm wherein implementation diversity drives improved verification success. Evaluated on a benchmark of 73 tasks, the method yields 154 verifiable programs after two rounds of repair, with at least one successfully verified variant for 67.1% of the tasks—substantially outperforming baseline approaches.
This work addresses the challenge of efficiently verifying the validity of statistical learning models on a given data distribution without trusting the learner. It introduces publicly verifiable certificates of statistical validity (pvCSVs), establishing the first non-interactive, publicly verifiable proof system for learning that applies to adaptive statistical query (SQ) algorithms. The proposed framework enables any user to verify a model’s performance on their own distribution with a sample complexity of only $O(\log k)$, a significant improvement over the $\tilde{O}(\sqrt{k})$ complexity of standard SQ learning algorithms. This result provides a systematic characterization of the capabilities and limitations of the SQ model in the context of verifiable learning.