🤖 AI Summary
This study addresses the challenges of large-scale proof obligations and the difficulty of balancing success rates against computational overhead in program verification by proposing the CoCo-Prover framework. This method formulates formal proving as cost-aware meta-level decision-making, employing a bi-level AND/OR hypergraph proof structure and lemma dependency graphs to enable symbolic topological selection. Furthermore, it treats expensive expert invocations as priced services for dynamic agent orchestration and routing. Evaluated in the Lean 4 environment across five benchmarks, CoCo-Prover achieves state-of-the-art solve rates—reaching up to 100%—while reducing computational costs by 30.9% compared to the strongest baseline. Ultimately, this work effectively unifies efficiency and economy in automated theorem proving.
📝 Abstract
Program verification establishes software correctness through machine-checkable proofs constructed in theorem provers. It's a guarantee especially valuable for code generated by large language models (LLMs), which is fluent but carries no assurance of correctness. Almost all existing provers, however, pursue pass rates alone at whatever sampling or search budget it takes, and overlook the success-vs-cost frontier; yet real software often carries hundreds of interdependent proof obligations, so what matters at scale is not whether one theorem can be proved, but how many can be proved economically. We introduce CoCo-Prover, which formalizes cost-efficient program proving as metalevel decision-making under cost, grounded on two-level proof graphs: an AND/OR proof hypergraph within each declaration is joined to a lemma-dependency graph across declarations; and at each step, it answers two questions: which open goals to select, and which actions to purchase on these goals. Selection stays symbolic as a topological pass over the proof graphs. Action choice is agent orchestration via metalevel decision-making: an agentic router treats every bounded specialist invocation as a separately priced, best-effort computation, matching heterogeneous specialist agents together with configurations, under evolved routing rules as evidence accumulates. On five program verification benchmarks in Lean 4 including function-level CLEVER, VERINA, and AlgoVeri, and repository-level NTP4VC and Vero, we show that CoCo-Prover achieves a better success-vs-cost frontier than baselines including frontier coding agents and state-of-the-art LLM-based provers: it achieves the best solve rate on every benchmark and up to 100% on two benchmarks. It also reduces cost by up to 30.9% compared to the strongest baseline with the strongest LLM in our evaluation.