Can LLMs Reason Like Automated Theorem Provers for Rust Verification? VCoT-Bench: Evaluating via Verification Chain of Thought

📅 2026-03-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing evaluation methods treat Rust program verification as a black box, relying solely on binary outcomes to assess whether large language models (LLMs) generate valid proof hints, thereby failing to capture their logical reasoning capabilities. This work proposes VCoT-Lift, a novel framework that, for the first time, lifts the low-level reasoning traces of automated theorem provers into human-readable, high-level “verification chains of thought.” To enable fine-grained assessment, we construct VCoT-Bench—a benchmark comprising 1,988 tasks—evaluating LLMs along three dimensions: robustness to missing proofs, type coverage, and positional sensitivity. Our evaluation of ten state-of-the-art LLMs reveals their fragile performance in formal Rust verification, demonstrating a significant gap between current models and the capabilities of dedicated theorem provers.

Technology Category

Knowledge Representation and Reasoning: Automated Reasoning and Theorem ProvingNatural Language Processing: Safety and RobustnessMachine Learning: Large Multimodal Models (LMMs)

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
📝 Abstract
As Large Language Models (LLMs) increasingly assist secure software development, their ability to meet the rigorous demands of Rust program verification remains unclear. Existing evaluations treat Rust verification as a black box, assessing models only by binary pass or fail outcomes for proof hints. This obscures whether models truly understand the logical deductions required for verifying nontrivial Rust code. To bridge this gap, we introduce VCoT-Lift, a framework that lifts low-level solver reasoning into high-level, human-readable verification steps. By exposing solver-level reasoning as an explicit Verification Chain-of-Thought, VCoT-Lift provides a concrete ground truth for fine-grained evaluation. Leveraging VCoT-Lift, we introduce VCoT-Bench, a comprehensive benchmark of 1,988 VCoT completion tasks for rigorously evaluating LLMs' understanding of the entire verification process. VCoT-Bench measures performance along three orthogonal dimensions: robustness to varying degrees of missing proofs, competence across different proof types, and sensitivity to the proof locations. Evaluation of ten state-of-the-art models reveals severe fragility, indicating that current LLMs fall well short of the reasoning capabilities exhibited by automated theorem provers.
Problem

Research questions and friction points this paper is trying to address.

LLM reasoning
Rust verification
automated theorem proving
verification chain of thought
program correctness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Verification Chain-of-Thought
VCoT-Lift
Rust verification
LLM reasoning evaluation
automated theorem proving
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zichen Xie
University of Virginia
W
Wenxi Wang
University of Virginia