Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of cross-scale generalization and repetitive retraining in reinforcement learning (RL)-based network defense agents by proposing a Planner-Executor hierarchical architecture built upon frozen large language models (LLMs). Leveraging the Cyberwheel high-fidelity simulation environment, this work systematically compares RL and LLM hybrid configurations, substituting conventional RL components with LLMs to enable zero-shot, retraining-free autonomous defense across varying network scales. The findings demonstrate that zero-shot LLMs exhibit scalable advantages from planning to execution within hierarchical defense frameworks, wherein tactical execution capability is critical for realizing their control efficacy. Notably, a 70B-parameter model reduces the attacker’s lateral movement success rate to approximately 1%, consistently maintaining robust defensive performance across multi-scale networks.
πŸ“ Abstract
An autonomous cyber defender trained with reinforcement learning (RL) is typically tied to the network on which it was trained, limiting its ability to generalize as network scale changes. Hierarchical RL reduces decision complexity by separating strategic targeting from tactical execution, but it does not eliminate this retraining dependence. We investigate whether frozen, zero-shot large language models (LLMs) can provide retraining-free control in hierarchical cyber defense and how performance changes as LLM control is extended from planning to execution. We formulate a controller-agnostic planner-executor hierarchy in which the planner selects a subnet to defend over a fixed horizon and the executor selects defensive actions within that subnet. Using the high fidelity Cyberwheel environment, with its built-in automated red team agent mapped to the MITRE ATT&CK framework, we compare RL+RL, LLM+RL, and LLM+LLM configurations using six models ranging from 3B to 70B parameters, including two cybersecurity-specialized models, across small, medium, and large networks. Replacing only the planner with an LLM yields limited gains as network size increases. In contrast, extending LLM control to execution produces notable improvements for sufficiently capable models. For instance, a frozen general purpose 70B model holds successful lateral movement to approximately 1% of steps and attacker impact near zero across all three network scales using the same model weights, while the RL baseline is retrained for each scale. Our results show that sufficiently capable frozen LLMs can maintain strong defensive performance across the evaluated network scales without task-specific retraining, while also indicating that strong tactical execution is important to realizing the benefits of LLM-based control.
Problem

Research questions and friction points this paper is trying to address.

Autonomous Cyber Defense
Hierarchical Reinforcement Learning
Large Language Models
Network Generalization
Zero-shot Control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Cyber Defense
Large Language Models
Zero-shot Generalization
Retraining-free Control
Reinforcement Learning
πŸ”Ž Similar Papers