VulContextBench: A Benchmark for Security Context Retrieval in Coding Agents

πŸ“… 2026-09-26
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation of existing vulnerability detection benchmarks, which evaluate only final verdicts and thus cannot distinguish data-flow reasoning from reliance on pretraining memorization. We construct a benchmark comprising 111 real-world vulnerability-introducing commits, manually audited to eliminate label noise and annotate gold-standard contexts. Uniquely, we propose scoring models based on retrieved evidence rather than final conclusions, introducing a multi-granularity evaluation framework at the file, chunk, and line levels. Experiments reveal that while state-of-the-art coding agents can explore most critical code, their final citation rates fall 37 to 73 percentage points below their viewing rates. This discrepancy highlights a significant gap between models’ ability to discover evidence and their capacity to explicitly cite it in their reasoning.
πŸ“ Abstract
Vulnerability-detection benchmarks score the verdict an agent reaches, not the evidence it gathered. A model that recalls a CVE from pretraining therefore scores the same as one that traced the data flow. We study a task where this difference matters, deciding whether a commit introduces a vulnerability. Instead of scoring the verdict, we score whether the agent retrieved the code its conclusion depends on. We present VulContextBench, a benchmark of 111 vulnerability-introducing commits (VICs) across 83 repositories, 63 CWEs, and five languages. Existing datasets label such commits by tracing a fix back through the version history, which often points to the wrong commit. We therefore audit every case by hand against an explicit four-criterion definition of a VIC, so the benchmark does not inherit that label noise. Each case is annotated with gold context, 464 code blocks in total, each tagged by its role in the evidence for the vulnerability. We evaluate seven frontier models with precision, recall and F1 at three granularities (file, block, and line), scored separately on the context an agent viewed while exploring and on the context it finally declared as evidence. The gap between the two is the main finding. Every model opens most of the gold context while exploring, but reports only part of it as evidence. At the level of code blocks, the share of the gold context a model reports is 37 to 73 percentage points below the share it viewed. Qwen3-Coder-Next views 86.3% of the lines in annotated code blocks but cites only 12.9% in its final report. GPT-5.5, which cites the most, views 73% and reports 36%. These results highlight a gap between finding relevant code and selecting it for the final report, which verdict-level benchmarks cannot reveal.
Problem

Research questions and friction points this paper is trying to address.

Vulnerability Detection
Context Retrieval
Coding Agents
Benchmark
Vulnerability-Introducing Commits
Innovation

Methods, ideas, or system contributions that make the work stand out.

Security Context Retrieval
Vulnerability-Introducing Commits
Benchmark Evaluation
Coding Agents
Evidence Grounding
πŸ”Ž Similar Papers
No similar papers found.
Yikun Li
Yikun Li
Postdoctoral Researcher
Artificial intelligenceSoftware EngineeringCyber Security
J
Jinfeng Jiang
Singapore Management University
Y
Yieh Yuheng
Singapore Management University
Ting Zhang
Ting Zhang
Monash University
Software EngineeringCyber SecurityInformation Retrieval
Y
Yide Yin
GovTech
L
Leow Wen Bin
GovTech
E
Eng Lieh Ouh
Singapore Management University
L
Lwin Khin Shar
Singapore Management University
David Lo
David Lo
Professor of Computer Science, Singapore Management University
AI4SESoftware AnalyticsSE4AISoftware MaintenanceSoftware Engineering