ReproBench: Benchmarking LLM Agents on Reproducing Vulnerability From Scratch

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing benchmarks that rely on preset environmental dependencies, rendering them inadequate for evaluating the ability of LLM agents to autonomously reproduce vulnerabilities from CVE identifiers. We propose the first zero-start, end-to-end vulnerability reproduction evaluation paradigm and an evidence-driven benchmark. This approach decouples the reproduction pipeline into six distinct stages and leverages 30 IoT firmware vulnerabilities, employing verifiable artifacts such as firmware images, unpacked binaries, and crash samples to enable fine-grained, independent assessment. Experimental results reveal that 45.3% of the runs exhibit simulated cheating, while only 5.3% achieve successful reproduction. Nevertheless, these findings demonstrate that frontier models possess the potential for fully autonomous vulnerability reproduction.
📝 Abstract
Large language model (LLM) agents are increasingly evaluated on cybersecurity tasks such as vulnerability reproduction, exploitation, and patching. However, existing cybersecurity benchmarks predominantly operate under a post-environment evaluation paradigm, i.e., handing the agent source code, a container, or an executable binary. This setup bypasses the critical environment reconstruction step, leaving a fundamental question for real-world vulnerability analysis: can an agent autonomously reconstruct the required execution environment and reproduce a vulnerability entirely from scratch? To address this gap, we present ReproBench, an evidence-grounded benchmark designed to evaluate agent capabilities in end-to-end vulnerability reproduction starting from solely a CVE identifier. ReproBench decomposes the full reproduction workflow into six distinct phases, and assesses performance on each phase independently using verifiable experimental artifacts: downloaded firmware images, unpacked binaries, granular analysis logs, and validated crash samples, among others. We instantiate ReproBench with 30 real-world IoT firmware vulnerabilities, which serve as ideal test cases for our from-scratch evaluation setting. Our evaluation demonstrates that 45.3% of test runs resort to vulnerability simulation - a prevalent remediation workaround adopted across all evaluated LLM agents - while only 5.3% of CVE-model pairs achieve successful reproduction of real-world vulnerabilities. Despite the low overall success rate, these non-trivial successful cases confirm that state-of-the-art LLM agents already possess the capacity for fully autonomous end-to-end vulnerability reproduction. Concurrently, our in-depth analysis of failed cases identifies core bottlenecks impeding LLM agents throughout the reproduction pipeline, offering actionable insights for subsequent research.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
vulnerability reproduction
benchmarking
environment reconstruction
cybersecurity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vulnerability Reproduction
LLM Agents
Benchmark
Environment Reconstruction
IoT Firmware
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Liang He
Institute of Software, Chinese Academy of Sciences; University of Chinese Academy of Sciences
S
Sheng Wu
Institute of Software, Chinese Academy of Sciences; University of Chinese Academy of Sciences
H
Haomiao Hao
Institute of Software, Chinese Academy of Sciences; University of Chinese Academy of Sciences
H
Hongduo Zhao
Institute of Software, Chinese Academy of Sciences; University of Chinese Academy of Sciences
Jia Yan
Jia Yan
The Hong Kong University of Science and Technology (Guangzhou)
Wireless Communications and NetworkingMobile Edge ComputingEdge Intelligence
Purui Su
Purui Su
Institute of Software, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Key Laboratory of System Software (Chinese Academy of Sciences)