Ask the Expert: LLM-Guided Reinforcement Learning for Autonomous Cyber Defense

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the low sample efficiency of reinforcement learning (RL) in cybersecurity defense and the high latency incurred by online deployment of large language models (LLMs). To this end, we propose a training-time guidance framework featuring an asymmetric design. During training, an LLM is employed solely to generate strategic suggestions that assist proximal policy optimization (PPO) via hierarchical reward shaping. For online deployment, the framework entirely eliminates LLM dependency, retaining only the pure RL policy. Experimental results demonstrate that the proposed approach significantly improves sample efficiency and outperforms existing baselines. Furthermore, it achieves efficient online defense with zero LLM overhead while preserving optimal terminal returns.
📝 Abstract
Policy-based reinforcement learning (RL) approaches have produced promising results for autonomous cyber defense; however, they are sample-inefficient in settings where defenders must respond under delayed, partial observations with actions from large action spaces. While large language models (LLMs) may reason semantically about security state space, high latency and trust assumptions prevent attractive in-line deployment models. We introduce Ask the Expert, a training-time guidance framework which first summarizes hard cyber-defense states, then intermittently queries an LLM for host-level defensive recommendations via a constrained action interface, and finally transforms those recommendations into tiered reward shaping for use with PPO. Because the LLM is discarded after training, deployment is a pure RL policy. Across TTCP CAGE CC1 and CC2 and both attacker types, this asymmetric design improves sample efficiency over PPO and outperforms the evaluated potential-based reward shaping (PBRS) baselines, while retaining the strongest terminal mean and requiring no LLM dependency at deployment time.
Problem

Research questions and friction points this paper is trying to address.

Autonomous Cyber Defense
Reinforcement Learning
Sample Efficiency
Large Language Models
Partial Observations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Large Language Models
Reward Shaping
Autonomous Cyber Defense
Sample Efficiency
🔎 Similar Papers
No similar papers found.