KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of fine-grained metrics in existing evaluations for large language models (LLMs) generating precise cybersecurity CLI commands, where syntax errors frequently cause execution failures. We construct a manual-driven benchmark based on Kali Linux and propose deterministic normalization alongside alias-aware evaluation mechanisms. Furthermore, we design a multi-stage verification pipeline integrating LLM-based validation, sandboxed execution, and human-in-the-loop review to enable precise training without runtime rewards. Experiments reveal that open-source models achieve less than 42% accuracy under unconstrained settings. However, reinforcement learning training on our proposed benchmark enables an 8B-parameter model to perform comparably to a 685B MoE model, effectively bridging the gap in fine-grained CLI capability assessment.
📝 Abstract
LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.
Problem

Research questions and friction points this paper is trying to address.

Cybersecurity
Tool Use
Command-Line Interface
Large Language Models
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cybersecurity Benchmark
NL-to-CLI Translation
Verifiable Rewards
Multi-stage Verification
Reinforcement Learning
🔎 Similar Papers
No similar papers found.