HoF-Bench: Rediscovering Real AI-Discovered CVEs Without Frontier Models

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether non-state-of-the-art large language models (LLMs) can reliably reproduce real AI-discovered CVE vulnerabilities without any prior knowledge of the flaws. To this end, we introduce HoF-Bench, a benchmark comprising 95 genuine AI-identified CVEs, and establish a rigorous blind-testing protocol that prohibits the use of CVE identifiers, patches, or vulnerability descriptions, requiring models to pinpoint root causes and impact pathways solely from raw source code. Within a unified evaluation framework, we assess ten open-source and commercial small-scale models using a methodology involving four rounds of repeated analysis, optional context generation, and a multi-stage traceable classification pipeline. Our results demonstrate that a minimal LLM analyzer—without leveraging any cutting-edge models—successfully reproduces 65 CVEs (68%), confirming the strong reproducibility potential of non-frontier models in specific contexts, while also highlighting C-language infrastructure code as a persistent challenge.
📝 Abstract
LLM-based analyzers have begun finding real vulnerabilities in mature open-source projects: AISLE's analyzer is credited with more than 280 CVEs across 78 projects, including OpenSSL, curl, and GnuTLS. We introduce HoF-Bench (named after AISLE's public Hall of Fame), a benchmark built from 95 of these public AI-discovered CVEs across eight repositories pinned at vulnerable commits. Analyzers receive source and target-file scope but not CVE identifiers, descriptions, fixes, or expected mechanisms; a detector-blinded frontier-model judge credits only findings that identify the same code path, root cause, attack condition, and impact. A deliberately minimal LLM-based analyzer rediscovers up to 65 of the 95 CVEs (68%) under this strict protocol. No frontier model performs detection anywhere in the study. The ten detector backbones are five open-weight models (21B--284B total parameters, 3--13B active) and five proprietary small or "flash"-tier models. All of them run in the fixed scaffold with four repeated passes, an optional generated-context stage, and a replayable multi-round triage stage (7,600 model--CVE pass records). Difficulty is strongly structured by language; the CVEs missed by every model concentrate in C infrastructure code. HoF-Bench provides a compact test bed for comparing vulnerability scanners, their reliability across repeated runs, and the candidate volume they create. The dataset is available at https://huggingface.co/datasets/aisleinc/HoF-Bench.
Problem

Research questions and friction points this paper is trying to address.

vulnerability detection
CVE rediscovery
LLM-based analyzers
benchmark
software security
Innovation

Methods, ideas, or system contributions that make the work stand out.

HoF-Bench
LLM-based vulnerability detection
detector-blinded evaluation
AI-discovered CVEs
open-weight models
🔎 Similar Papers
No similar papers found.
Petr Simecek
Petr Simecek
CEITEC Masaryk University
Large Language ModelsDeep LearningBenchmarksGenetics & GenomicsMedia Monitoring
E
Elnaz Babayeva
AISLE, San Francisco, CA, USA and Prague, Czech Republic
J
Jiri Balhar
AISLE, San Francisco, CA, USA and Prague, Czech Republic
M
Michal Bida
AISLE, San Francisco, CA, USA and Prague, Czech Republic
M
Michal Buran
AISLE, San Francisco, CA, USA and Prague, Czech Republic
V
Vaclav Cadek
AISLE, San Francisco, CA, USA and Prague, Czech Republic
L
Luigino Camastra
AISLE, San Francisco, CA, USA and Prague, Czech Republic
T
Tomas Dulka
AISLE, San Francisco, CA, USA and Prague, Czech Republic
M
Michal Janocko
AISLE, San Francisco, CA, USA and Prague, Czech Republic
T
Tomas Klohna
AISLE, San Francisco, CA, USA and Prague, Czech Republic
P
Pavel Kohout
AISLE, San Francisco, CA, USA and Prague, Czech Republic
O
Ondrej Kokes
AISLE, San Francisco, CA, USA and Prague, Czech Republic
A
Adam Krivka
AISLE, San Francisco, CA, USA and Prague, Czech Republic
J
Jakub Kubik
AISLE, San Francisco, CA, USA and Prague, Czech Republic
P
Patrik Mada
AISLE, San Francisco, CA, USA and Prague, Czech Republic
I
Igor Morgenstern
AISLE, San Francisco, CA, USA and Prague, Czech Republic
M
Marek Pavelka
AISLE, San Francisco, CA, USA and Prague, Czech Republic
J
Joshua Rogers
AISLE, San Francisco, CA, USA and Prague, Czech Republic
P
Petr Stastny
AISLE, San Francisco, CA, USA and Prague, Czech Republic
J
Jan Tattermusch
AISLE, San Francisco, CA, USA and Prague, Czech Republic
Dmitrijs Trizna
Dmitrijs Trizna
Microsoft Corporation
artificial intelligencecybersecurity
M
Martin Votruba
AISLE, San Francisco, CA, USA and Prague, Czech Republic
G
Guido Vranken
AISLE, San Francisco, CA, USA and Prague, Czech Republic
J
Jakub Zikl
AISLE, San Francisco, CA, USA and Prague, Czech Republic
E
Evelina Gabasova
AISLE, San Francisco, CA, USA and Prague, Czech Republic