hybrid trace verification

Design and implement verifiers that analyze stepwise execution or reasoning traces by applying deterministic, rule-based checks to mechanizable steps, resolving dependencies and satisfying cross-step constraints. Where semantics cannot be fully mechanized, integrate targeted LLM audits in a deterministic-plus-LLM hybrid verifier, support verifier-addressable checking, and localize and repair step-level errors to produce stepwise verification and repair outputs.

hybridtraceverification

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.21
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

The 4/$delta$ Bound: Designing Predictable LLM-Verifier Systems for Formal Method Guarantee

Nov 30, 2025
PD
Pierre Dantas
🏛️ The University of Manchester | Federal University of Amazonas

Existing LLM–verifier collaboration frameworks for formal verification lack theoretical guarantees, leading to unstable behavior such as non-termination or divergence. Method: We propose the first formally verified LLM–verifier framework with provable termination and convergence: we model the interaction as a discrete-time Markov chain, establish a quantitative relationship between error-reduction probability δ and expected iteration count, and derive a convergence theorem yielding an analytical upper bound of 4/δ on expected iterations. Contribution/Results: This enables systematic, predictability-driven system design—replacing heuristic tuning with rigorous resource planning. Empirical evaluation across >90,000 tasks demonstrates universal convergence, with measured convergence factor (C_f approx 1.0), confirming tight alignment between theory and practice. The framework provides a quantifiable foundation for resource allocation in safety-critical software verification.

Develops a formal framework with provable guarantees for LLM-verifier convergence and terminationEstablishes design thresholds and predictable performance zones for safety-critical software verificationModels LLM-verifier interaction as a Markov chain using error-reduction probability to bound expected iterations

Autonomous multimodal LLM agents pose escalating risks of loss of control, necessitating rigorous behavioral oversight. Method: We propose a verifiability-centered control architecture featuring runtime cryptographic signing and symbolic action attestation; a lightweight audit agent with challenge-response protocols; and OPERA, a novel benchmark shifting evaluation from “prevention” to “detection–response.” Our approach integrates cryptographic signatures, symbolic reasoning, and constraint-logic verification, rigorously validated via red-teaming and robustness testing against prompt and persona manipulation. Contribution/Results: Experiments demonstrate significant improvements in detection latency and attribution reliability for covert misalignment behaviors, while maintaining high observability and strong auditability across diverse adversarial scenarios.

Detecting and remedying misalignment between agent intent and behaviorEnsuring controllability and auditability in autonomous LLM-based agentsMeasuring resilience of verifiability mechanisms against adversarial strategies

LLMs as verification oracles for Solidity

Sep 23, 2025
MB
Massimo Bartoletti
🏛️ University of Cagliari

Business logic errors in smart contracts are a leading cause of substantial financial losses, yet existing formal verification tools suffer from high learning barriers and limited expressiveness of specification languages. Method: This paper presents the first systematic evaluation of reasoning-capable large language models (e.g., GPT-5) as oracles for Solidity contract verification, proposing a novel “AI + formal methods” hybrid paradigm. We design a mixed quantitative–qualitative evaluation framework, benchmarking LLM outputs against industrial-grade tools (e.g., SolCMC, Certora) on real-world audit tasks. Results: Empirical evaluation demonstrates that GPT-5 effectively detects complex logical vulnerabilities in realistic auditing scenarios, achieving performance comparable to specialized formal verifiers. Our core contribution is establishing LLMs as lightweight, scalable verification oracles—overcoming key usability bottlenecks of traditional formal methods—and thereby opening a new, practical pathway for enhancing smart contract security.

Addressing limitations of formal verification tools for business logic errorsAssessing LLM capability to reason about arbitrary contract-specific propertiesEvaluating LLMs as verification oracles for smart contract correctness

This work addresses the security risks posed by large language model (LLM) agents during tool invocation, such as inadvertent leakage of sensitive data or overwriting of critical records—hazards for which existing approaches lack verifiable guarantees. To bridge this gap, the paper introduces a novel integration of System-Theoretic Process Analysis (STPA) with formal specifications to systematically identify hazards in agent workflows and derive enforceable safety requirements. These requirements are then translated into executable constraints on data flows and tool invocation sequences. Building upon an enhanced Model Context Protocol (MCP) framework, the approach incorporates structured capability control and trust-labeling mechanisms to enable proactive, verifiable protection of tool interactions. By significantly reducing reliance on manual verification, this method advances LLM agent design from empirical reliability toward a paradigm grounded in formal security assurances.

enterprise riskhazard mitigationLLM agents

Large language model inference lacks output determinism due to floating-point non-associativity, dynamic batching, and varying GPU reduction orders. This work proposes a scheduling-based speculative validation mechanism that introduces speculative execution into deterministic inference for the first time. By employing lightweight validate-and-rollback cycles combined with fixed-shape reduction scheduling, the approach incurs overhead only for requests requiring determinism, while remaining compatible with dynamic batching and requiring minimal modification to existing GPU kernels. The method decouples determinism guarantees from low-level implementation details, achieving high throughput and significantly outperforming baseline strategies such as disabling dynamic batching or rewriting kernel functions.

dynamic batchingfloating-point non-associativityGPU kernels

Latest Papers

What's happening recently
View more

This study addresses the high specification burden of static verification tools like VeriFast, which, despite their ability to verify complex heap-manipulating programs using separation logic, require costly manual annotation. For the first time, this work systematically evaluates the capability of large language models (LLMs) to automatically generate C function specifications for VeriFast, examining ten prominent LLMs, eight prompting strategies, and three input formats through both quantitative and qualitative analyses across two experimental phases. Results show that LLM-generated specifications achieve over 91% functional behavioral consistency and a 31.4% verification success rate; notably, 94% of failures stem from insufficient domain-specific knowledge of VeriFast. The findings highlight domain adaptation as a critical bottleneck and propose effective strategies to improve verification success rates.

LLMseparation logicspecification generation

This work addresses the challenge of statically verifying semantic consistency between natural language business requirements and their code implementations. It proposes a two-stage, runtime-free approach: first leveraging large language models to extract structured rules from requirements while identifying ambiguous or contradictory statements, and then performing static code auditing based on this intermediate representation. By integrating natural language processing with static analysis, the method mitigates hallucination and context loss in large models through rule structuring, enabling requirement-aware early validation. Evaluated on an automotive cybersecurity case study, the approach successfully detects semantic deviations, offers a novel solution to the test oracle problem, and significantly enhances left-shifted verification capabilities.

business logic validationcode compliancenatural-language requirements

Large language models often introduce subtle, hard-to-detect bugs when generating complex software, compromising reliability. This work proposes the first fully automated, project-level code generation and verification framework based on an interactive theorem prover (ITP). The approach separates code with side effects into C++ while formalizing pure logical components in the ITP Rocq, where they are automatically verified and extracted for integration. When proofs fail, the concrete counterexample states guide an LLM agent to autonomously repair the code. In experiments, the system generated 1,859 lines of verified Rocq code and extracted 2,848 lines of C++ within 30 minutes, passing 265 unit tests and 12 hours of AFL++ fuzzing with zero crashes or hangs—outperforming Dafny’s backend, which failed to complete verification under identical conditions.

Formal VerificationInteractive Theorem ProvingLarge-scale Code Generation

This work addresses the significant disparity in verifiability among semantically equivalent yet structurally diverse programs, a key bottleneck in generating high-assurance software. The authors propose Diversify2Verify, a novel approach that leverages large language models to synthesize diverse recursive and imperative implementations of the same task, integrates the Why3 platform for automatic contract inference and formal verification, and introduces a verifier-guided annotation repair mechanism to enhance verifiability. This study is the first to systematically expose the verifiability gap across equivalent program variants and establishes a new paradigm wherein implementation diversity drives improved verification success. Evaluated on a benchmark of 73 tasks, the method yields 154 verifiable programs after two rounds of repair, with at least one successfully verified variant for 67.1% of the tasks—substantially outperforming baseline approaches.

automated verifiabilityimplementation diversityprogram verification

This work addresses the challenge of silent error propagation in multi-step reasoning, where early logical mistakes or hallucinations often lead large language models to produce confidently incorrect conclusions. The authors propose a zero-shot verification and repair framework that formalizes natural language reasoning traces into a structured domain-specific language (DSL), explicitly encoding step dependencies, executable quantitative expressions, and deductive structures. By integrating deterministic checks—such as computational correctness and constraint satisfaction—with semantic auditing via large language models, the approach enables training-free, step-level error detection and correction. Evaluated on mathematical reasoning, robotic planning, and kinship inference tasks, the method substantially outperforms zero-shot baselines, significantly enhancing the reasoning accuracy of mainstream large models without requiring domain-specific data or exemplars.

Chain-of-Thoughthallucinationlogical error

Hot Scholars

BF

Bernd Finkbeiner

Professor of Computer Science, CISPA Helmholtz Center for Information Security
Reactive SystemsVerificationSynthesisTemporal Logic
AF

Angelo Ferrando

Assistant Professor at the University of Modena and Reggio Emilia
Artificial intelligenceFormal MethodsRuntime VerificationMulti-Agent Systems
RB

Raven Beutner

CISPA Helmholtz Center for Information Security
Formal Methods
PG

Pranshav Gajjar

North Carolina State University
Deep LearningApplied Artificial IntelligenceImage ProcessingLarge Language Models
BB

Boris Bellalta

Professor. Wireless Networking Group, Dept. of Information and Communication Technologies, UPF
Wireless NetworksWi-Fi / 802.11Performance Evaluation