Correct Verdicts, Flawed Reasoning: Structured Auditing of LLM-based Vulnerability Reasoning

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of fabricated reasoning, unverifiability, and the failure of conventional evaluation methods in LLM vulnerability analysis by proposing the VERA framework. VERA replaces free-text outputs with Structured Reasoning Records (SRR) and integrates chain-of-thought prompting with deterministic rule checking to enable multi-stage auditing. Furthermore, it introduces an automated mutation testing benchmark to ensure verifiability. Experimental results demonstrate that 60% of correct judgments are accompanied by covert logical errors, and VERA successfully identifies 87% of the reasoning flaws overlooked by traditional methods. Ultimately, this work renders the security analysis process for large language models both quantifiable and auditable.
📝 Abstract
Large Language Models (LLMs) are increasingly deployed for automated software vulnerability analysis. Binary classification alone is insufficient; practitioners need explanations to triage bugs and engineer patches. Standard practice relies on Chain-of-Thought (CoT) prompting, but free-form reasoning allows models to obscure logical leaps, hallucinated execution steps, and internal inconsistencies behind plausible prose. Our manual audit reveals that approximately 60% of correct vulnerability verdicts are accompanied by fabricated or unverifiable claims, and free-form explanations allow reasoning errors to evade LLM-as-a-judge evaluation. We present Vulnerability Explanation Reasoning Auditor (VERA), an automated framework for auditing LLM vulnerability reasoning. Rather than accepting free-form text, VERA asks models to output a Structured Reasoning Record (SRR) encoding tracked pointers, memory operations, and state transitions in machine-readable fields. A multi-stage judge audits each SRR against eight reasoning failure modes using deterministic checks, with LLM calls reserved for semantic interpretation. The standardized SRR schema also enables automated mutation testing to benchmark judges at scale without human annotation. Our evaluation shows reasoning flaws occur in correct verdicts just as frequently as incorrect ones, and VERA exposes 87% of reasoning errors that free-form LLM-as-judge systematically miss.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Vulnerability Reasoning
Chain-of-Thought Hallucination
Automated Auditing
Software Vulnerability Analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Structured Reasoning Record
Vulnerability Reasoning Auditing
Mutation Testing
Chain-of-Thought
LLM-as-a-Judge
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Boyue Caroline Hu
Boyue Caroline Hu
University of Toronto
K
Kaivalya Ahir
Carnegie Mellon University, USA
R
Ronghao Ni
Carnegie Mellon University, USA
Limin Jia
Limin Jia
Carnegie Mellon University
Programming LanguagesFormal MethodsSecurity