Score
Designs, builds, or analyzes specifications, proofs, tests, and other verification artifacts that establish whether a program, algorithm, model, or system satisfies its intended functional and non‑functional properties. Work includes creating formal specifications and correctness proofs, devising test suites and runtime checks, and diagnosing counterexamples or failures to determine compliance with requirements.
This work addresses the limited adoption of formal verification, which often requires expert-written annotations such as preconditions, postconditions, and loop invariants. To overcome this barrier, the authors propose a novel approach that leverages large language models (LLMs) in conjunction with assertions from test cases as static oracles to automatically generate Dafny verification annotations from code annotated with natural language comments. The method features an iterative refinement process guided by verifier feedback over multiple rounds and uniquely integrates multi-model LLM collaboration with a closed-loop verifier feedback mechanism. A VS Code plugin was developed to support practical deployment. Evaluated on 110 Dafny programs, the approach achieves a 98.2% annotation correctness rate within at most eight repair iterations. Empirical results highlight that proof-assistant-style annotation remains a key challenge for LLMs, while user feedback on the plugin was notably positive.
Existing code-level formal verification tools scale poorly to large-scale software, while mainstream unit-level verification relies heavily on manual effort, often missing critical defects. This paper proposes the “Unit Proof Framework” research agenda—the first systematic definition of a unit verification paradigm supporting automated decoupling and independent verification of code units. Methodologically, it integrates formal verification, program analysis, modular verification, and automated toolchain design, with deep alignment to industrial development practices (e.g., AWS workflows). Its core contributions include: (1) establishing a scalable, engineering-friendly unit verification methodology; (2) characterizing a taxonomy of key technical challenges; (3) overcoming bottlenecks inherent in manual verification; and (4) significantly improving early detection of code-level defects. Collectively, this work lays the theoretical foundation and provides a practical technical pathway for building high-assurance, deployable automated verification infrastructure.
This paper addresses hyperproperties—higher-order system requirements encompassing information-flow security, knowledge reasoning, and robustness, which span multiple execution traces—by proposing the first unified logical and algorithmic framework covering the entire verification lifecycle. Methodologically, it rigorously characterizes the expressive power and decidability boundaries of classical temporal logics (LTL, CTL, S1S) over hyperproperties; then introduces a novel multi-trace synchronization modeling and quantifier alternation handling mechanism grounded in higher-order temporal logic, constraint solving, and symbolic automata. Key contributions include: (i) a comprehensive taxonomy and complexity-theoretic characterization of hyperproperty logics; (ii) an open-source verification toolchain supporting HyperLTL and HyperCTL*; and (iii) end-to-end support for core verification tasks—including satisfiability checking, model checking, runtime monitoring, and controller synthesis.
Automatically verifying C programs generated by large language models (LLMs) remains challenging due to their syntactic and semantic irregularities, which hinder formal verification. Method: This paper proposes SynVer—a novel framework that tightly integrates LLM-based program synthesis with formal verification. SynVer introduces verifiability-aware biasing mechanisms operating at both syntactic and semantic levels to guide LLMs toward generating verification-friendly code. It further incorporates separation logic (SL) specifications and the Verified Software Toolchain (VST) to enable end-to-end, fully automated verification—from specification to C implementation to machine-checked safety proofs. Results: Evaluated on diverse benchmarks covering basic coding tasks, SL assertions, and API specifications, SynVer significantly improves the automatic verification success rate of LLM-generated C programs. Empirical results demonstrate its scalability, robustness, and effectiveness in bridging the gap between neural code generation and rigorous formal assurance.
Automated verification of interactive console I/O programs in Haskell education remains challenging due to the dynamic, history-dependent nature of student implementations. Method: We propose a lightweight, formal behavioral specification language that uniquely integrates global state and execution history, expressed via regex-like syntax; its trace-based semantics enable probabilistic testing and scalable verification through *sampleable validity*. Contribution/Results: Our system automatically validates student submissions against behavioral specifications and supports pedagogical closed-loop applications—including real-time feedback generation, example solution synthesis, and exercise randomization. Empirical evaluation demonstrates substantial improvements in test coverage and pedagogical adaptability while preserving formal rigor. To our knowledge, this is the first framework for verifying interactive behaviors in functional programming education that simultaneously achieves theoretical soundness and practical deployability.
Existing large code models struggle to generate executable intermediate formal specifications, limiting precise verification and repair of program behavioral errors. This work proposes SpecCoder, a novel framework that focuses on generating executable inline assertions at critical program locations, thereby transforming static annotations into verifiable evidence. SpecCoder employs verification-guided training, fine-tuning the Qwen2.5-Coder series models using correct programs, behavioral mutants, and multi-round specification refinement trajectories. Evaluated on the HumanExec benchmark, SpecCoder substantially improves the correctness (+55.8%), completeness (+358.1%), and assertion validity (+26.6%) of inline specifications, significantly enhancing program verification and repair capabilities.
This work addresses the high cost of manually writing formal specifications and the limitations of existing large language model (LLM)-based approaches that require white-box access to source code, thereby posing intellectual property and deployment constraints. The authors propose a black-box-driven method that leverages only test code and dynamic execution traces to generate candidate Java Modeling Language (JML) specifications via an LLM. These candidates are locally validated using bounded model checking, and an iterative feedback loop refines them based on verification outcomes. This approach is the first to enable fully automated formal specification generation without any access to the program’s internal structure. Evaluated on the SpecGenBench benchmark, it demonstrates that test-derived information effectively guides specification synthesis, while also highlighting critical challenges in checker compatibility and diagnostic feedback, substantially enhancing industrial applicability.
This work addresses the significant disparity in verifiability among semantically equivalent yet structurally diverse programs, a key bottleneck in generating high-assurance software. The authors propose Diversify2Verify, a novel approach that leverages large language models to synthesize diverse recursive and imperative implementations of the same task, integrates the Why3 platform for automatic contract inference and formal verification, and introduces a verifier-guided annotation repair mechanism to enhance verifiability. This study is the first to systematically expose the verifiability gap across equivalent program variants and establishes a new paradigm wherein implementation diversity drives improved verification success. Evaluated on a benchmark of 73 tasks, the method yields 154 verifiable programs after two rounds of repair, with at least one successfully verified variant for 67.1% of the tasks—substantially outperforming baseline approaches.
This study addresses the challenges of applying formal verification to production-grade software, where high modeling costs and consistency risks in fault handling hinder adoption. By integrating runtime execution traces with formal specifications, the authors verify a real-world restaurant point-of-sale (POS) payment workflow and leverage large language models (LLMs) to automatically generate these specifications. Their analysis reveals that the structural form—not the natural language phrasing—of specifications primarily governs LLM-generated correctness, and uncovers a shared “relevant oracle failure” issue between code and simulators. Extending fault models to include crash-recovery, stale reads, and retries, the team conducts simulation-based audits, verifying core protocol correctness, identifying and reproducing seven fault-handling vulnerabilities, and revalidating after fixes. They also expose how deviations in API response structures render recovery paths unreachable—a finding consistently replicated across seven LLMs from two vendors.
This work addresses the challenge that counterexamples generated by formal verification often consist of numerous low-level Boolean variables, rendering them difficult for developers to interpret at the application-domain level. To bridge this gap, the paper proposes a novel hierarchical explanation method that integrates predicate relevance metrics with dependency graph analysis—a first-time fusion of these two techniques—to automatically extract human-readable, domain-oriented explanations from logical formulas. By leveraging formal modeling and a dedicated explanation-generation algorithm, the approach produces concise and semantically clear descriptions of failure causes across multiple case studies. Empirical results demonstrate that the method significantly outperforms existing techniques, offering effective support for fault localization in practical verification tasks.