🤖 AI Summary
Existing LLM–verifier collaboration frameworks for formal verification lack theoretical guarantees, leading to unstable behavior such as non-termination or divergence.
Method: We propose the first formally verified LLM–verifier framework with provable termination and convergence: we model the interaction as a discrete-time Markov chain, establish a quantitative relationship between error-reduction probability δ and expected iteration count, and derive a convergence theorem yielding an analytical upper bound of 4/δ on expected iterations.
Contribution/Results: This enables systematic, predictability-driven system design—replacing heuristic tuning with rigorous resource planning. Empirical evaluation across >90,000 tasks demonstrates universal convergence, with measured convergence factor (C_f approx 1.0), confirming tight alignment between theory and practice. The framework provides a quantifiable foundation for resource allocation in safety-critical software verification.
📝 Abstract
The idea of using Formal Verification tools with large language models (LLMs) has enabled scaling software verification beyond manual workflows. However, current methods remain unreliable. Without a solid theoretical footing, the refinement process can wander; sometimes it settles, sometimes it loops back, and sometimes it breaks away from any stable trajectory. This work bridges this critical gap by developing an LLM-Verifier Convergence Theorem, providing the first formal framework with provable guarantees for termination and convergence. We model the interaction between the LLM and the verifier as a discrete-time Markov Chain, with state transitions determined by a key parameter: the error-reduction probability ($delta$). The procedure reaching the Verified state almost surely demonstrates that the program terminates for any $delta>0$, with an expected iteration count bounded by $mathbb{E}[n] leq 4/delta$. We then stress-tested this prediction in an extensive empirical campaign comprising more than 90,000 trials. The empirical results match the theory with striking consistency. Every single run reached verification, and the convergence factor clustered tightly around $C_fapprox$ 1.0. Consequently, the bound mirrors the system's actual behavior. The evidence is sufficiently robust to support dividing the workflow into three distinct operating zones: marginal, practical, and high-performance. Consequently, we establish the design thresholds with absolute confidence. Together, the theoretical guarantee and the experimental evidence provide a clearer architectural foundation for LLM-assisted verification. Heuristic tuning no longer has to be carried out by the system. Engineers gain a framework that supports predictable resource planning and performance budgeting, precisely what is needed before deploying these pipelines into safety-critical software environments.