Cheap to Hypothesize, Costly to Verify: The Defense Surface of Agentic Vulnerability Discovery

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the asymmetry in LLM agent-based vulnerability discovery, where hypothesis generation is inexpensive yet verification is costly. We propose a novel defense strategy leveraging certifiable security decoys. By integrating CVE chains with an adaptive mechanism based on unreachable bridging, and employing reachability analysis alongside computational complexity theory, our approach decouples private efficient verification from public intractability. This effectively redirects agents' verification resources away from genuine vulnerabilities toward decoys. Evaluated across 33 OSS-Fuzz projects, the proposed method consumes approximately 50% of the verification budget, reducing the true vulnerability discovery rate by 38.7% to 60.4%. Furthermore, it maintains robustness even under informed adversarial conditions.
📝 Abstract
Autonomous LLM agents turn vulnerability discovery into a repository-scale search: they generate many vulnerability hypotheses but can verify only a subset under a finite budget. We show that autonomous vulnerability discovery exhibits a hypothesis-verification asymmetry, where verifying a candidate hypothesis through reachability analysis, execution, and proof-of-concept construction is substantially more expensive than forming it. Under a finite resource budget, this makes autonomous discovery a resource-bounded selective-verification process, further exposing verification effort as a unique defense surface. We present RedHerring, which inserts certifiably safe decoys that divert verification effort from real vulnerabilities. Each decoy combines a CVE-derived vulnerability chain that attracts verification with a false bridge that keeps its dangerous sink unreachable. A private certificate lets the defender verify this property efficiently, while establishing the same fact from the released repository requires solving a computationally hard problem. RedHerring further adapts each decoy to the target repository so that it reads as ordinary program logic. Across 33 OSS-Fuzz projects, 70 evaluation instances, and five models under matched budgets, RedHerring reduces real vulnerabilities discovered by 38.7-60.4%. Trajectory analysis shows that agents spend 30.6-51.5% of completion tokens and an estimated 32.5-49.9% of runtime verifying decoys, showing that RedHerring redirects a substantial fraction of the fixed search budget toward decoys. When explicitly informed that decoys may be present, the agent adapts its search strategy, yet RedHerring still reduces vulnerabilities discovered by 37.2% relative to an informed Baseline, showing that its effectiveness does not depend on decoy secrecy.
Problem

Research questions and friction points this paper is trying to address.

autonomous vulnerability discovery
LLM agents
hypothesis-verification asymmetry
resource-bounded search
AI security
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Vulnerability Discovery
Hypothesis-Verification Asymmetry
Decoy-based Defense
Resource-bounded Search
Certifiable Safety
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kaikai Zhang
The Hong Kong University of Science and Technology
Zihan Zhang
Zihan Zhang
Southern University of Science and Technology
HCI
Yuchong Xie
Yuchong Xie
HKUST
Security
Zesen Liu
Zesen Liu
Ph.D. Student, HKUST
Security
S
Shuangjie Yao
The Hong Kong University of Science and Technology
Z
Zhixiang Zhang
The Hong Kong University of Science and Technology
Dongdong She
Dongdong She
Hong Kong University of Science and Technology
SecurityMachine LearningProgram AnalysisFuzzing