Score
Design and build methods and pipelines that synthesize program invariants (including loop invariants) by combining symbolic inference and neural or LLM-based components; concretely, this involves generating candidate invariants symbolically, integrating or reconciling those symbolic invariants with neural/LLM outputs, and producing deterministic invariants that seed automated verification. Practitioners implement hybrid or purely symbolic workflows that can operate with or without LLM calls and that scale verification runs on local machines.
This work addresses the automated synthesis of loop invariants for formal verification of looping programs. We propose a generate-and-verify closed-loop framework that tightly integrates large language models (LLMs)—specifically reasoning-optimized variants such as OpenAI O1, O1-mini, and O3-mini—with the Z3 SMT solver. Invariant synthesis proceeds via counterexample-guided iterative refinement, where verification feedback directly drives LLM inference optimization. To our knowledge, this is the first approach achieving deep, symbol-level synergy between LLMs and SMT solvers in formal verification. Evaluated on the Code2Inv benchmark (133 benchmarks), our method achieves 100% coverage—surpassing the prior state-of-the-art (107/133)—with an average of only 1–2 LLM invocations and runtime of 14–55 seconds per benchmark. The results demonstrate substantial improvements in precision, efficiency, and generalization, validating the substantive potential of LLMs in formal deductive reasoning.
This work presents the first systematic evaluation of large language models’ (LLMs) ability to infer and repair program loop invariants without auxiliary information. We adopt an empirical framework encompassing diverse open- and closed-source LLMs across multiple scales, integrating domain-knowledge augmentation and few-shot prompting to quantify performance on standard benchmarks for inductive invariant generation and logical defect repair. Results show that LLMs achieve up to 78% success in invariant generation but only 16% in invariant repair—revealing a critical bottleneck in deep logical correction. A key contribution is the identification of auxiliary information—particularly loop semantics prompts and correct examples—as decisive for improving repair accuracy. Our study establishes a reproducible evaluation paradigm for LLM-driven automated program safety analysis and provides concrete, actionable pathways for enhancing invariant repair capabilities.
Traditional loop invariant generation tools exhibit limited precision and applicability on real-world programs where complex data structures intertwine with intricate control flow. Method: This paper proposes ACInv, the first static-analysis-driven, LLM-augmented framework for invariant synthesis. It extracts loop semantic features to construct structured prompts for LLM-based candidate invariant generation, and introduces an LLM-powered semantic evaluator that dynamically refines candidates via strengthening, weakening, or rejection. Contribution/Results: ACInv is the first approach to support template-level invariant generation for user-defined data structures. Evaluated on benchmarks containing complex data structures, ACInv achieves a 21% higher overall solving rate than AutoSpec, while matching its performance on purely numeric programs. Moreover, the generated invariants are reusable and significantly improve practicality for industrial-scale program verification.
Neural-symbolic program verification suffers from a semantic gap—termed the “embedding gap”—between neural components and symbolic logic. Method: This paper formally defines the embedding gap and proposes an end-to-end formal verification framework comprising: (1) a domain-specific language (DSL) to declaratively specify problem-space properties; (2) a multi-backend compiler that enables declarative, compilable mapping from the problem space to the embedding space, seamlessly interfacing PyTorch, Marabou, and Lean; and (3) modular, co-verification across training environments, neural verifiers, and theorem provers. Contribution/Results: We demonstrate fully automated, reproducible, and mathematically rigorous safety verification on a simplified autonomous driving system equipped with a neural controller. Our approach systematically bridges the semantic divide between neural and symbolic verification, enabling principled integration of learning-based and logic-based reasoning within a unified formal framework.
Automatically verifying C programs generated by large language models (LLMs) remains challenging due to their syntactic and semantic irregularities, which hinder formal verification. Method: This paper proposes SynVer—a novel framework that tightly integrates LLM-based program synthesis with formal verification. SynVer introduces verifiability-aware biasing mechanisms operating at both syntactic and semantic levels to guide LLMs toward generating verification-friendly code. It further incorporates separation logic (SL) specifications and the Verified Software Toolchain (VST) to enable end-to-end, fully automated verification—from specification to C implementation to machine-checked safety proofs. Results: Evaluated on diverse benchmarks covering basic coding tasks, SL assertions, and API specifications, SynVer significantly improves the automatic verification success rate of LLM-generated C programs. Empirical results demonstrate its scalability, robustness, and effectiveness in bridging the gap between neural code generation and rigorous formal assurance.
This work addresses the challenge of automatically inferring loop invariants in programs with multiple interacting loops. It proposes a novel neurosymbolic framework that uniquely integrates obligation-guided reasoning with weakest precondition refinement, explicitly modeling inter-loop dependencies through loop-level abstractions and propagating proof obligations accordingly. By synergistically combining large language models with formal verification techniques, the approach introduces a deductive feedback mechanism to iteratively refine candidate invariants. Evaluated on a new benchmark comprising classical algorithms, the method successfully solves 72 out of 82 multi-loop problems, substantially outperforming existing approaches, while also maintaining state-of-the-art performance on single-loop tasks.
This work addresses the challenge of verification failures in loop invariant synthesis caused by local reasoning errors in large language models (LLMs). To this end, the authors propose LORIS, a novel framework that integrates formal verification of natural-language reasoning steps with feedback-driven iterative refinement. LORIS automatically translates LLM-generated natural language invariants into first-order logic and employs formal verification to detect logical inconsistencies, which are then used to generate targeted feedback for guiding the model to correct its reasoning trajectory. Experimental results demonstrate that LORIS achieves a 93.1% success rate on a benchmark of 460 C programs and exhibits strong robustness on 50 challenging programs involving nonlinear properties, substantially enhancing the reliability of LLM-based reasoning in program verification.
This work addresses the challenging problem of loop invariant synthesis, which is inherently undecidable and difficult for existing learning-based methods to solve due to their inability to generate complete and ordered sequences of invariants. The paper proposes an incremental ICE framework that, for the first time, integrates the incremental reasoning principle from IC3 into learning-based invariant inference. By introducing lemma-specific learning objectives and a counterexample filtering mechanism, the approach leverages large language models (LLMs) to produce ordered lemma sequences, with ICE-DT serving as a fallback. This enables lemma-level controllable learning that effectively combines LLMs with symbolic reasoning. Evaluated on 367 linear and 50 nonlinear benchmarks, the method solves 349 and 47 instances respectively, with average runtimes of 15.2 and 8.8 seconds—outperforming state-of-the-art LLM baselines by solving 12–24% more instances and achieving 36–63% speedups, while significantly surpassing strong non-LLM baselines.
This work addresses the bottleneck of loop invariant synthesis in formal verification by introducing VerIbmc, the first fully local neurosymbolic framework that operates without reliance on cloud-based large language model APIs. The approach integrates deterministic symbolic reasoning, locally deployed open-source large language models (ranging from 7B to 120B parameters), the ESBMC model checker, and a structured feedback mechanism, supporting both Chain-of-Thought and Tree-of-Thought prompting strategies. Evaluated on 499 benchmark problems, the best configuration (GPT-OSS-120B) solves 431 instances (86.4%), matching the performance of state-of-the-art cloud-based tools. Notably, the symbolic component alone solves 75 problems and substantially enhances the efficacy of weaker models, achieving efficient verification while preserving code privacy and minimizing computational cost.
This work addresses the challenges of low reliability, poor auditability, high cost, and security risks associated with runtime invocation of large language models (LLMs) in high-stakes enterprise workflows. To overcome these issues, we propose a novel “compiled AI” paradigm, wherein LLMs generate executable code during compilation, eliminating the need for model calls at runtime and thereby ensuring deterministic execution. We present the first systematic application of this paradigm to high-risk scenarios, integrating constrained code generation, a four-stage verification pipeline, template-embedded business logic functions, and an operation-oriented evaluation framework to jointly achieve reliability, auditability, and security. Experiments demonstrate a 96% success rate on function-calling tasks with zero runtime token consumption; 80.0% and 80.4% accuracy on critical field extraction and line-item recognition in document intelligence tasks; and strong security performance, with 96.7% prompt injection detection accuracy and 87.5% static analysis precision without false positives.