🤖 AI Summary
This paper addresses the lack of standardized, rigorous evaluation benchmarks for large language models (LLMs) in adversarial cybersecurity applications—particularly penetration testing. We systematically review 15 prototype systems and their empirical testing practices drawn from 16 scholarly works. Through comparative analysis of testbed architectures, evaluation metric frameworks, and security experimentation methodologies, we identify three pervasive deficiencies: poor reproducibility, weak real-world representativeness (especially the limited transferability of CTF-based scenarios to actual red-teaming), and narrow assessment scope (lacking established baselines, qualitative analysis, and human-factor considerations). To address these gaps, we propose the first dedicated evaluation framework for LLM-driven adversarial security, featuring an extensible testbed, standardized baseline construction, a hybrid quantitative–qualitative metric suite, and human-in-the-loop assessment dimensions. Based on this framework, we derive seven actionable best-practice recommendations—establishing a scientifically grounded, robust, and operationally viable evaluation paradigm for the field.
📝 Abstract
Large Language Models (LLMs) have emerged as a powerful approach for driving offensive penetration-testing tooling. This paper analyzes the methodology and benchmarking practices used for evaluating Large Language Model (LLM)-driven attacks, focusing on offensive uses of LLMs in cybersecurity. We review 16 research papers detailing 15 prototypes and their respective testbeds. We detail our findings and provide actionable recommendations for future research, emphasizing the importance of extending existing testbeds, creating baselines, and including comprehensive metrics and qualitative analysis. We also note the distinction between security research and practice, suggesting that CTF-based challenges may not fully represent real-world penetration testing scenarios.