🤖 AI Summary
This study investigates whether non-state-of-the-art large language models (LLMs) can reliably reproduce real AI-discovered CVE vulnerabilities without any prior knowledge of the flaws. To this end, we introduce HoF-Bench, a benchmark comprising 95 genuine AI-identified CVEs, and establish a rigorous blind-testing protocol that prohibits the use of CVE identifiers, patches, or vulnerability descriptions, requiring models to pinpoint root causes and impact pathways solely from raw source code. Within a unified evaluation framework, we assess ten open-source and commercial small-scale models using a methodology involving four rounds of repeated analysis, optional context generation, and a multi-stage traceable classification pipeline. Our results demonstrate that a minimal LLM analyzer—without leveraging any cutting-edge models—successfully reproduces 65 CVEs (68%), confirming the strong reproducibility potential of non-frontier models in specific contexts, while also highlighting C-language infrastructure code as a persistent challenge.
📝 Abstract
LLM-based analyzers have begun finding real vulnerabilities in mature open-source projects: AISLE's analyzer is credited with more than 280 CVEs across 78 projects, including OpenSSL, curl, and GnuTLS. We introduce HoF-Bench (named after AISLE's public Hall of Fame), a benchmark built from 95 of these public AI-discovered CVEs across eight repositories pinned at vulnerable commits. Analyzers receive source and target-file scope but not CVE identifiers, descriptions, fixes, or expected mechanisms; a detector-blinded frontier-model judge credits only findings that identify the same code path, root cause, attack condition, and impact. A deliberately minimal LLM-based analyzer rediscovers up to 65 of the 95 CVEs (68%) under this strict protocol. No frontier model performs detection anywhere in the study. The ten detector backbones are five open-weight models (21B--284B total parameters, 3--13B active) and five proprietary small or "flash"-tier models. All of them run in the fixed scaffold with four repeated passes, an optional generated-context stage, and a replayable multi-round triage stage (7,600 model--CVE pass records). Difficulty is strongly structured by language; the CVEs missed by every model concentrate in C infrastructure code. HoF-Bench provides a compact test bed for comparing vulnerability scanners, their reliability across repeated runs, and the candidate volume they create. The dataset is available at https://huggingface.co/datasets/aisleinc/HoF-Bench.