Score
Design and implement verifiers that analyze stepwise execution or reasoning traces by applying deterministic, rule-based checks to mechanizable steps, resolving dependencies and satisfying cross-step constraints. Where semantics cannot be fully mechanized, integrate targeted LLM audits in a deterministic-plus-LLM hybrid verifier, support verifier-addressable checking, and localize and repair step-level errors to produce stepwise verification and repair outputs.
Existing LLM–verifier collaboration frameworks for formal verification lack theoretical guarantees, leading to unstable behavior such as non-termination or divergence. Method: We propose the first formally verified LLM–verifier framework with provable termination and convergence: we model the interaction as a discrete-time Markov chain, establish a quantitative relationship between error-reduction probability δ and expected iteration count, and derive a convergence theorem yielding an analytical upper bound of 4/δ on expected iterations. Contribution/Results: This enables systematic, predictability-driven system design—replacing heuristic tuning with rigorous resource planning. Empirical evaluation across >90,000 tasks demonstrates universal convergence, with measured convergence factor (C_f approx 1.0), confirming tight alignment between theory and practice. The framework provides a quantifiable foundation for resource allocation in safety-critical software verification.
Autonomous multimodal LLM agents pose escalating risks of loss of control, necessitating rigorous behavioral oversight. Method: We propose a verifiability-centered control architecture featuring runtime cryptographic signing and symbolic action attestation; a lightweight audit agent with challenge-response protocols; and OPERA, a novel benchmark shifting evaluation from “prevention” to “detection–response.” Our approach integrates cryptographic signatures, symbolic reasoning, and constraint-logic verification, rigorously validated via red-teaming and robustness testing against prompt and persona manipulation. Contribution/Results: Experiments demonstrate significant improvements in detection latency and attribution reliability for covert misalignment behaviors, while maintaining high observability and strong auditability across diverse adversarial scenarios.
Business logic errors in smart contracts are a leading cause of substantial financial losses, yet existing formal verification tools suffer from high learning barriers and limited expressiveness of specification languages. Method: This paper presents the first systematic evaluation of reasoning-capable large language models (e.g., GPT-5) as oracles for Solidity contract verification, proposing a novel “AI + formal methods” hybrid paradigm. We design a mixed quantitative–qualitative evaluation framework, benchmarking LLM outputs against industrial-grade tools (e.g., SolCMC, Certora) on real-world audit tasks. Results: Empirical evaluation demonstrates that GPT-5 effectively detects complex logical vulnerabilities in realistic auditing scenarios, achieving performance comparable to specialized formal verifiers. Our core contribution is establishing LLMs as lightweight, scalable verification oracles—overcoming key usability bottlenecks of traditional formal methods—and thereby opening a new, practical pathway for enhancing smart contract security.
This work addresses the security risks posed by large language model (LLM) agents during tool invocation, such as inadvertent leakage of sensitive data or overwriting of critical records—hazards for which existing approaches lack verifiable guarantees. To bridge this gap, the paper introduces a novel integration of System-Theoretic Process Analysis (STPA) with formal specifications to systematically identify hazards in agent workflows and derive enforceable safety requirements. These requirements are then translated into executable constraints on data flows and tool invocation sequences. Building upon an enhanced Model Context Protocol (MCP) framework, the approach incorporates structured capability control and trust-labeling mechanisms to enable proactive, verifiable protection of tool interactions. By significantly reducing reliance on manual verification, this method advances LLM agent design from empirical reliability toward a paradigm grounded in formal security assurances.
Large language model inference lacks output determinism due to floating-point non-associativity, dynamic batching, and varying GPU reduction orders. This work proposes a scheduling-based speculative validation mechanism that introduces speculative execution into deterministic inference for the first time. By employing lightweight validate-and-rollback cycles combined with fixed-shape reduction scheduling, the approach incurs overhead only for requests requiring determinism, while remaining compatible with dynamic batching and requiring minimal modification to existing GPU kernels. The method decouples determinism guarantees from low-level implementation details, achieving high throughput and significantly outperforming baseline strategies such as disabling dynamic batching or rewriting kernel functions.
This study addresses the high specification burden of static verification tools like VeriFast, which, despite their ability to verify complex heap-manipulating programs using separation logic, require costly manual annotation. For the first time, this work systematically evaluates the capability of large language models (LLMs) to automatically generate C function specifications for VeriFast, examining ten prominent LLMs, eight prompting strategies, and three input formats through both quantitative and qualitative analyses across two experimental phases. Results show that LLM-generated specifications achieve over 91% functional behavioral consistency and a 31.4% verification success rate; notably, 94% of failures stem from insufficient domain-specific knowledge of VeriFast. The findings highlight domain adaptation as a critical bottleneck and propose effective strategies to improve verification success rates.
This work addresses the challenge of statically verifying semantic consistency between natural language business requirements and their code implementations. It proposes a two-stage, runtime-free approach: first leveraging large language models to extract structured rules from requirements while identifying ambiguous or contradictory statements, and then performing static code auditing based on this intermediate representation. By integrating natural language processing with static analysis, the method mitigates hallucination and context loss in large models through rule structuring, enabling requirement-aware early validation. Evaluated on an automotive cybersecurity case study, the approach successfully detects semantic deviations, offers a novel solution to the test oracle problem, and significantly enhances left-shifted verification capabilities.
Large language models often introduce subtle, hard-to-detect bugs when generating complex software, compromising reliability. This work proposes the first fully automated, project-level code generation and verification framework based on an interactive theorem prover (ITP). The approach separates code with side effects into C++ while formalizing pure logical components in the ITP Rocq, where they are automatically verified and extracted for integration. When proofs fail, the concrete counterexample states guide an LLM agent to autonomously repair the code. In experiments, the system generated 1,859 lines of verified Rocq code and extracted 2,848 lines of C++ within 30 minutes, passing 265 unit tests and 12 hours of AFL++ fuzzing with zero crashes or hangs—outperforming Dafny’s backend, which failed to complete verification under identical conditions.
This work addresses the significant disparity in verifiability among semantically equivalent yet structurally diverse programs, a key bottleneck in generating high-assurance software. The authors propose Diversify2Verify, a novel approach that leverages large language models to synthesize diverse recursive and imperative implementations of the same task, integrates the Why3 platform for automatic contract inference and formal verification, and introduces a verifier-guided annotation repair mechanism to enhance verifiability. This study is the first to systematically expose the verifiability gap across equivalent program variants and establishes a new paradigm wherein implementation diversity drives improved verification success. Evaluated on a benchmark of 73 tasks, the method yields 154 verifiable programs after two rounds of repair, with at least one successfully verified variant for 67.1% of the tasks—substantially outperforming baseline approaches.
This work addresses the challenge of silent error propagation in multi-step reasoning, where early logical mistakes or hallucinations often lead large language models to produce confidently incorrect conclusions. The authors propose a zero-shot verification and repair framework that formalizes natural language reasoning traces into a structured domain-specific language (DSL), explicitly encoding step dependencies, executable quantitative expressions, and deductive structures. By integrating deterministic checks—such as computational correctness and constraint satisfaction—with semantic auditing via large language models, the approach enables training-free, step-level error detection and correction. Evaluated on mathematical reasoning, robotic planning, and kinship inference tasks, the method substantially outperforms zero-shot baselines, significantly enhancing the reasoning accuracy of mainstream large models without requiring domain-specific data or exemplars.