🤖 AI Summary
This study addresses the limited and fragmented coverage of existing evaluation benchmarks for attacks on large language models, which impedes comprehensive assessment of defensive capabilities. To bridge this gap, the authors propose the first 4×6 objective-technique matrix grounded in the STRIDE threat modeling framework, synthesizing 507 distinct inference-time attack types extracted from a systematic review of 932 publications. This effort yields a scalable attack taxonomy and a benchmark coverage auditing framework. Through systematic literature synthesis, attack clustering, and benchmark mapping, the analysis reveals that prevailing benchmarks cover at most 25% of the defined threat surface, with several high-severity attack categories entirely absent. The project is open-sourced, providing a repository of 2,521 attack groups to enable ongoing community-driven auditing and evolutionary tracking of evaluation benchmarks.
📝 Abstract
We introduce a reusable framework for auditing whether LLM attack benchmarks collectively cover the threat surface: a 4$\times$6 Target $\times$ Technique matrix grounded in STRIDE, constructed from a 507-leaf taxonomy -- 401 data-populated and 106 threat-model-derived leaves -- of inference-time attacks extracted from 932 arXiv security studies (2023--2026). The matrix enables benchmark-external validation -- auditing collective coverage rather than individual benchmark consistency. Applying it to six public benchmarks reveals that the three primary frameworks (HarmBench, InjecAgent, AgentDojo) occupy non-overlapping cells covering at most 25\% of the matrix, while entire STRIDE threat categories (Service Disruption, Model Internals) lack any standardized evaluation, despite published attacks in these categories achieving 46$\times$ token amplification and 96\% attack success rates through mechanisms which no benchmark tests. The corpus of 2,521 unique attack groups further reveals pervasive naming fragmentation (up to 29 surface forms for a single attack) and heavy concentration in Safety \& Alignment Bypass, structural properties invisible at smaller scale. The taxonomy, attack records, and coverage mappings are released as extensible artifacts; as new benchmarks emerge, they can be mapped onto the same matrix, enabling the community to track whether evaluation gaps are closing.