Towards an Automated Test of LLM Security Knowledge

📅 2026-07-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing safety knowledge evaluations for large language models (LLMs), which rely on costly, manually curated benchmarks that are difficult to scale. The authors propose a partially automated approach that leverages authoritative information from consumer protection agencies (CPAs) as an external knowledge source. By analyzing response consistency across two critical safety topics—identity theft and impersonation scams—the method automatically identifies knowledge gaps in LLMs. Integrating publicly available CPA data with natural language understanding and consistency analysis, the approach is validated on five prominent models from the Gemini and GPT families. Results demonstrate its ability to reliably differentiate models based on their safety knowledge, offering a scalable and practical framework for ongoing LLM safety assessment.
📝 Abstract
Large language models (LLMs) are increasingly used for a range of software, hardware and human-centered security tasks. Consequently, LLM performance on security tasks is an active area of measurement and research, often with a focus on identifying areas in which LLM security ``knowledge'' may be insufficient. Popular strategies for identifying LLM security knowledge gaps include building corpora of challenge questions or task benchmarks, strategies that require substantial manual work and security expertise to design and execute. We introduce a partially-automated method for assessing LLM knowledge of a security area.The method uses authoritative information from Consumer Protection Agencies (CPAs) to identify instability in LLM responses that can be indicative of knowledge gaps. We demonstrate the method for 2 security topics, identity theft and impostor scams, and 5 LLMs in 2 leading LLM families, Gemini and GPT, using publicly available information about identity theft and impostor scams from 6 CPAs.The method distinguishes between models that have and don't have sufficient knowledge to accurately identify the security topics in text narratives.
Problem

Research questions and friction points this paper is trying to address.

LLM security knowledge
automated evaluation
knowledge gaps
security benchmarks
Consumer Protection Agencies
Innovation

Methods, ideas, or system contributions that make the work stand out.

automated security evaluation
LLM knowledge gap
Consumer Protection Agencies
response instability
security benchmarking
🔎 Similar Papers
No similar papers found.