The 4/$delta$ Bound: Designing Predictable LLM-Verifier Systems for Formal Method Guarantee

📅 2025-11-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing LLM–verifier collaboration frameworks for formal verification lack theoretical guarantees, leading to unstable behavior such as non-termination or divergence. Method: We propose the first formally verified LLM–verifier framework with provable termination and convergence: we model the interaction as a discrete-time Markov chain, establish a quantitative relationship between error-reduction probability δ and expected iteration count, and derive a convergence theorem yielding an analytical upper bound of 4/δ on expected iterations. Contribution/Results: This enables systematic, predictability-driven system design—replacing heuristic tuning with rigorous resource planning. Empirical evaluation across >90,000 tasks demonstrates universal convergence, with measured convergence factor (C_f approx 1.0), confirming tight alignment between theory and practice. The framework provides a quantifiable foundation for resource allocation in safety-critical software verification.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Safety and RobustnessConstraint Satisfaction and Optimization: Satisfiability Modulo Theories

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
📝 Abstract
The idea of using Formal Verification tools with large language models (LLMs) has enabled scaling software verification beyond manual workflows. However, current methods remain unreliable. Without a solid theoretical footing, the refinement process can wander; sometimes it settles, sometimes it loops back, and sometimes it breaks away from any stable trajectory. This work bridges this critical gap by developing an LLM-Verifier Convergence Theorem, providing the first formal framework with provable guarantees for termination and convergence. We model the interaction between the LLM and the verifier as a discrete-time Markov Chain, with state transitions determined by a key parameter: the error-reduction probability ($delta$). The procedure reaching the Verified state almost surely demonstrates that the program terminates for any $delta>0$, with an expected iteration count bounded by $mathbb{E}[n] leq 4/delta$. We then stress-tested this prediction in an extensive empirical campaign comprising more than 90,000 trials. The empirical results match the theory with striking consistency. Every single run reached verification, and the convergence factor clustered tightly around $C_fapprox$ 1.0. Consequently, the bound mirrors the system's actual behavior. The evidence is sufficiently robust to support dividing the workflow into three distinct operating zones: marginal, practical, and high-performance. Consequently, we establish the design thresholds with absolute confidence. Together, the theoretical guarantee and the experimental evidence provide a clearer architectural foundation for LLM-assisted verification. Heuristic tuning no longer has to be carried out by the system. Engineers gain a framework that supports predictable resource planning and performance budgeting, precisely what is needed before deploying these pipelines into safety-critical software environments.
Problem

Research questions and friction points this paper is trying to address.

Develops a formal framework with provable guarantees for LLM-verifier convergence and termination
Models LLM-verifier interaction as a Markov chain using error-reduction probability to bound expected iterations
Establishes design thresholds and predictable performance zones for safety-critical software verification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Developed LLM-Verifier Convergence Theorem for formal guarantees
Modeled interaction as Markov Chain with error-reduction probability parameter
Established 4/δ bound for predictable iteration count and termination
🔎 Similar Papers
No similar papers found.
P
Pierre Dantas
Dept. of Computer Science, The University of Manchester, UK
L
Lucas Cordeiro
Dept. of Computer Science, The University of Manchester, UK
Youcheng Sun
Youcheng Sun
MBZUAI, Honorary SL@UoM
Trustworthy AIAutomated Reasoning
W
Waldir Junior
Dept. of Electrical Engineering, Federal University of Amazonas (UFAM), Brazil