Talk is (Not) Cheap: A Taxonomy and Benchmark Coverage Audit for LLM Attacks

📅 2026-05-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited and fragmented coverage of existing evaluation benchmarks for attacks on large language models, which impedes comprehensive assessment of defensive capabilities. To bridge this gap, the authors propose the first 4×6 objective-technique matrix grounded in the STRIDE threat modeling framework, synthesizing 507 distinct inference-time attack types extracted from a systematic review of 932 publications. This effort yields a scalable attack taxonomy and a benchmark coverage auditing framework. Through systematic literature synthesis, attack clustering, and benchmark mapping, the analysis reveals that prevailing benchmarks cover at most 25% of the defined threat surface, with several high-severity attack categories entirely absent. The project is open-sourced, providing a repository of 2,521 attack groups to enable ongoing community-driven auditing and evolutionary tracking of evaluation benchmarks.
📝 Abstract
We introduce a reusable framework for auditing whether LLM attack benchmarks collectively cover the threat surface: a 4$\times$6 Target $\times$ Technique matrix grounded in STRIDE, constructed from a 507-leaf taxonomy -- 401 data-populated and 106 threat-model-derived leaves -- of inference-time attacks extracted from 932 arXiv security studies (2023--2026). The matrix enables benchmark-external validation -- auditing collective coverage rather than individual benchmark consistency. Applying it to six public benchmarks reveals that the three primary frameworks (HarmBench, InjecAgent, AgentDojo) occupy non-overlapping cells covering at most 25\% of the matrix, while entire STRIDE threat categories (Service Disruption, Model Internals) lack any standardized evaluation, despite published attacks in these categories achieving 46$\times$ token amplification and 96\% attack success rates through mechanisms which no benchmark tests. The corpus of 2,521 unique attack groups further reveals pervasive naming fragmentation (up to 29 surface forms for a single attack) and heavy concentration in Safety \& Alignment Bypass, structural properties invisible at smaller scale. The taxonomy, attack records, and coverage mappings are released as extensible artifacts; as new benchmarks emerge, they can be mapped onto the same matrix, enabling the community to track whether evaluation gaps are closing.
Problem

Research questions and friction points this paper is trying to address.

LLM attacks
benchmark coverage
threat surface
evaluation gaps
attack taxonomy
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM security benchmarking
STRIDE-based threat taxonomy
coverage audit framework
attack surface mapping
benchmark standardization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Karthik Raghu Iyer
Palo Alto Networks
Y
Yazdan Jamshidi
Palo Alto Networks
N
Nicholas Bray
Palo Alto Networks
A
Alexey A. Shvets
Palo Alto Networks