Does AI Help Cyber Attackers or Defenders? Evidence from Nonpublic Vulnerabilities and Subsequent Attacks

πŸ“… 2026-10-05
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation of existing public benchmarks in distinguishing AI models’ memorization of known vulnerabilities from their genuine security capabilities against unknown ones. To this end, it proposes an evaluation framework comprising five non-public, unpatched vulnerability environments to systematically assess AI performance in exploitation, remediation, and sustained adversarial defense. The methodology employs human-verified deterministic scorers rather than LLM-based judges to conduct automated security benchmarking across both open-source and proprietary models. The findings reveal significant cross-environment variability in remediation capabilities and document 92 successful post-defense breaches following initial mitigation, thereby demonstrating the inadequacy of single-shot static evaluations. Ultimately, this work establishes a more reliable dynamic paradigm for assessing AI code security.
πŸ“ Abstract
The release decision for frontier AI systems increasingly relies on cyber capability benchmarks, yet public vulnerability benchmarks can expose agents to previously published advisories, exploits, and fixes, making it difficult to distinguish prior exposure from capability on unseen vulnerabilities. We evaluate open-weight and proprietary AI models on exploit generation, vulnerability repair and subsequent attacks in five nonpublic software environments, including vulnerabilities we privately disclosed while they remained unpatched. Researcher-developed and reviewed deterministic graders, not LLM judges, determine task scores. Comparisons with 209 disclosed vulnerabilities and cryptographic challenges reveal substantial variation across systems and vulnerability types. Repair scores exceed attack scores in two nonpublic environments and fall below them in three. Passing an initial security test is also insufficient: another exploit succeeds in 92 of 524 non-independent defender test intervals after the initial exploit is stopped. These results motivate vulnerability-specific attack-repair comparisons and subsequent resistance tests.
Problem

Research questions and friction points this paper is trying to address.

AI cybersecurity capability
nonpublic vulnerabilities
exploit generation
vulnerability repair
security benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

nonpublic vulnerabilities
deterministic graders
exploit generation
vulnerability repair
subsequent resistance tests
πŸ”Ž Similar Papers
No similar papers found.