VulContextBench: A Benchmark for Security Context Retrieval in Coding Agents
This study addresses the limitation of existing vulnerability detection benchmarks, which evaluate only final verdicts and thus cannot distinguish data-flow reasoning from reliance on pretraining memorization. We construct a benchmark comprising 111 real-world vulnerability-introducing commits, manually audited to eliminate label noise and annotate gold-standard contexts. Uniquely, we propose scoring models based on retrieved evidence rather than final conclusions, introducing a multi-granularity evaluation framework at the file, chunk, and line levels. Experiments reveal that while state-of-the-art coding agents can explore most critical code, their final citation rates fall 37 to 73 percentage points below their viewing rates. This discrepancy highlights a significant gap between models’ ability to discover evidence and their capacity to explicitly cite it in their reasoning.