Score
Design and build end-to-end verification procedures and tooling that specify formal correctness properties and implement, compose, and tune automated verifier checks and program-based verifiers. Analyze and refine those workflows by integrating manual/human review and safety-verification steps, defining schemas and validation rules, and generating or reconstructing realistic edit sequences and tests to validate outputs and measure verifier coverage and error modes.
Automatically generating high-quality, correct, and complete formal specifications—such as those in JML—remains a significant challenge: existing approaches often produce specifications that pass syntactic validation yet suffer from semantic inaccuracies or insufficient coverage. This work proposes VeriAct, a novel framework that introduces Spec-Harness, the first evaluation mechanism capable of precisely assessing both correctness and completeness of generated specifications. VeriAct further establishes the first verification-guided agent system, leveraging large language models within a closed-loop iterative process that integrates code execution, formal verification, and feedback signals to collaboratively synthesize and repair specifications. Experimental results demonstrate that VeriAct substantially outperforms current methods on two benchmarks, yielding specifications that not only satisfy verifiers but also achieve higher standards of semantic correctness and completeness.
Natural language requirements are ill-suited for direct use in formal verification. Method: This work proposes an automated property generation framework integrating large language models (LLMs) with formal verification tools. It introduces an assertion generation mechanism extending beyond Linear Temporal Logic (LTL) to support numerical constraints and compositional system behavior modeling; integrates Claude 3.5 Sonnet with the ESBMC bounded model checker; and employs human-in-the-loop supervision to calibrate output quality. Contribution/Results: We first identify and characterize systematic impacts of LLM-induced model connection errors and numerical approximations on verification outcomes—reducing false positives and uncovering previously overlooked falsifiable scenarios. Evaluated on nine cyber-physical systems from Lockheed Martin, our approach achieves 46.5% verification accuracy—on par with NASA’s CoCoSim—while substantially lowering the barrier to formal verification and enhancing defect detection capability.
This work addresses the limited adoption of formal verification, which often requires expert-written annotations such as preconditions, postconditions, and loop invariants. To overcome this barrier, the authors propose a novel approach that leverages large language models (LLMs) in conjunction with assertions from test cases as static oracles to automatically generate Dafny verification annotations from code annotated with natural language comments. The method features an iterative refinement process guided by verifier feedback over multiple rounds and uniquely integrates multi-model LLM collaboration with a closed-loop verifier feedback mechanism. A VS Code plugin was developed to support practical deployment. Evaluated on 110 Dafny programs, the approach achieves a 98.2% annotation correctness rate within at most eight repair iterations. Empirical results highlight that proof-assistant-style annotation remains a key challenge for LLMs, while user feedback on the plugin was notably positive.
This work addresses the challenge that program specifications generated from natural language are often too weak or overly restrictive for effective verification, and existing approaches are constrained by a single verification paradigm, struggling to balance automation with expressiveness. The paper proposes Velvet, a multimodal verifier architecture that unifies dynamic testing, automated reasoning, and interactive proving within a certified program synthesis pipeline, enabling specification validation, task decomposition, and proof delegation. Built upon Lean, Velvet integrates random property-based testing, verification-condition-guided divide-and-conquer synthesis, and state-of-the-art AI-powered theorem provers. Experiments demonstrate that the approach effectively uncovers flaws in existing specifications on standard benchmarks, substantially increases the rate of fully verified solutions, and maintains consistent performance across different large language model backends.
Despite its efficacy in isolated projects, deductive verification has yet to achieve broad industrial adoption. To identify root barriers and key enablers, this paper conducts semi-structured interviews with 30 practitioners, followed by thematic analysis. We systematically uncover fundamental obstacles—including high proof maintenance overhead, limited automation, poor tool usability, and lack of workflow integration—as well as critical enabling factors. Diverging from prior work, we empirically establish *usability* and *workflow adaptability* as core dimensions governing adoption. Based on these findings, we propose three actionable improvement principles: (1) enhancing automation support for proof construction and evolution; (2) reducing proof maintenance burden through modularization and abstraction; and (3) deepening integration with IDEs and CI/CD pipelines. Our empirically grounded insights provide concrete, evidence-based guidance for tool developers, practitioners, and researchers—bridging the gap between academic verification techniques and engineering practice.
This work addresses the challenge of ensuring program correctness in natural language-to-code generation, which is often hindered by the absence of high-quality formal specifications. The authors propose VeriSpecGen, a framework that decomposes natural language requirements into atomic clauses through a traceable refinement mechanism, generates requirement-driven tests with explicit traceability mappings, and synthesizes formal specifications aligned with user intent by localizing and repairing faulty clauses upon verification failure. Integrating large language models (e.g., Claude Opus 4.5) with the Lean proof assistant, the approach leverages refinement trajectories to generate 343K training samples, substantially enhancing model generalization and reasoning capabilities. Evaluated on the VERINA SpecGen benchmark, VeriSpecGen achieves an accuracy of 86.6%, outperforming the best baseline by up to 31.8 percentage points and demonstrating a relative improvement of 62–106% in specification synthesis performance.
This study addresses the limitations of existing SysML verification approaches, which are often tool-dependent and restricted to performance properties, lacking support for automated validation of behavioral and interface requirements. To overcome these shortcomings, this work proposes a tool-agnostic, automated verification workflow driven by SysML test cases, integrating UML Testing Profile and behavioral diagram constructs to enable unified validation of multidimensional attributes—including behavior, timing, and state responses. The methodology was developed through a mixed-methods research strategy combining literature review and stakeholder interviews, and its efficacy was empirically validated across two independent SysML toolchains. The approach not only transcends the constraints of conventional parametric methods but also enables automatic traceability of verification results back to the original model elements.
This work addresses the lack of semantic guarantees in existing ladder diagram verification tools, which often leads to false negatives or false positives in safety violation detection due to imprecise translation into model checker inputs. To remedy this, the authors present the first K Framework–based, standards-faithful, and reusable executable formal semantics for IEC 61131-3 ladder diagrams. This semantics uniformly yields both an interpreter and a deductive verifier, serving as an independent audit benchmark for differential testing of translation processes. It accurately models contacts, coils, timers, counters, and retentive scan cycles, with machine-checked correctness verified via kprove. Applying this approach uncovered two real-world flaws in ESBMC: unsound certification of unsafe programs and generation of spurious counterexamples. Furthermore, it formally guarantees input/output behavioral equivalence with the standard for both combinational and latching logic.
This work addresses the significant disparity in verifiability among semantically equivalent yet structurally diverse programs, a key bottleneck in generating high-assurance software. The authors propose Diversify2Verify, a novel approach that leverages large language models to synthesize diverse recursive and imperative implementations of the same task, integrates the Why3 platform for automatic contract inference and formal verification, and introduces a verifier-guided annotation repair mechanism to enhance verifiability. This study is the first to systematically expose the verifiability gap across equivalent program variants and establishes a new paradigm wherein implementation diversity drives improved verification success. Evaluated on a benchmark of 73 tasks, the method yields 154 verifiable programs after two rounds of repair, with at least one successfully verified variant for 67.1% of the tasks—substantially outperforming baseline approaches.
Automatically generating accurate and verifiable ACSL (ANSI/ISO C Specification Language) specifications remains a key challenge in C program verification. This work presents the first systematic evaluation of the verifiability and stability of ACSL annotations produced by rule-based scripts, the Frama-C RTE plugin, and three large language models—DeepSeek-V3.2, GPT-5.2, and OLMo 3.1 32B Instruct—under a fully automatic, non-interactive, and learning-free setting. Using a unified verification framework based on the Frama-C WP plugin coupled with multiple SMT solvers, the study provides new empirical evidence on the quality of automated ACSL generation, solver sensitivity, and proof stability, thereby revealing both the strengths and limitations of current approaches.