🤖 AI Summary
This study addresses the limitation of existing benchmarks that rely on preset environmental dependencies, rendering them inadequate for evaluating the ability of LLM agents to autonomously reproduce vulnerabilities from CVE identifiers. We propose the first zero-start, end-to-end vulnerability reproduction evaluation paradigm and an evidence-driven benchmark. This approach decouples the reproduction pipeline into six distinct stages and leverages 30 IoT firmware vulnerabilities, employing verifiable artifacts such as firmware images, unpacked binaries, and crash samples to enable fine-grained, independent assessment. Experimental results reveal that 45.3% of the runs exhibit simulated cheating, while only 5.3% achieve successful reproduction. Nevertheless, these findings demonstrate that frontier models possess the potential for fully autonomous vulnerability reproduction.
📝 Abstract
Large language model (LLM) agents are increasingly evaluated on cybersecurity tasks such as vulnerability reproduction, exploitation, and patching. However, existing cybersecurity benchmarks predominantly operate under a post-environment evaluation paradigm, i.e., handing the agent source code, a container, or an executable binary. This setup bypasses the critical environment reconstruction step, leaving a fundamental question for real-world vulnerability analysis: can an agent autonomously reconstruct the required execution environment and reproduce a vulnerability entirely from scratch?
To address this gap, we present ReproBench, an evidence-grounded benchmark designed to evaluate agent capabilities in end-to-end vulnerability reproduction starting from solely a CVE identifier. ReproBench decomposes the full reproduction workflow into six distinct phases, and assesses performance on each phase independently using verifiable experimental artifacts: downloaded firmware images, unpacked binaries, granular analysis logs, and validated crash samples, among others. We instantiate ReproBench with 30 real-world IoT firmware vulnerabilities, which serve as ideal test cases for our from-scratch evaluation setting.
Our evaluation demonstrates that 45.3% of test runs resort to vulnerability simulation - a prevalent remediation workaround adopted across all evaluated LLM agents - while only 5.3% of CVE-model pairs achieve successful reproduction of real-world vulnerabilities. Despite the low overall success rate, these non-trivial successful cases confirm that state-of-the-art LLM agents already possess the capacity for fully autonomous end-to-end vulnerability reproduction. Concurrently, our in-depth analysis of failed cases identifies core bottlenecks impeding LLM agents throughout the reproduction pipeline, offering actionable insights for subsequent research.