🤖 AI Summary
This study addresses the low sample efficiency and high environment interaction costs of model-free deep reinforcement learning in autonomous network defense by introducing the world model paradigm to cybersecurity for the first time. Building upon the Dreamer framework, the proposed method constructs a world model that integrates vector, graph, and text representations through multimodal learning to capture latent network dynamics. Defense policies are optimized via imagined rollouts, with graph-structured representations identified as the core design for enhancing robustness and scalability. Experimental results demonstrate that this approach surpasses baseline performance using only thousands of interaction steps—saving millions of steps compared to PPO—while maintaining high efficiency as network scale expands.
📝 Abstract
Deep reinforcement learning has become a prominent approach to autonomous cyber defense. Existing methods are predominantly model-free and consequently require extensive environment interaction. World models provide an alternative by learning predictive dynamics and optimizing policies through imagined trajectories, yielding substantial gains in sample efficiency in robotics and embodied control. Extending this paradigm to cybersecurity raises a fundamental question: what should constitute the"world"in a cyber world model? We introduce CyberWorld, a Dreamer-style world modeling framework that learns latent cyber dynamics from vector, graph, textual, and multimodal representations of the defended network. Across all four scoreable CyberWheel attack strategies, the graph-based CyberWorld variant exceeds a strategy-agnostic control after 3.6k-15.8k environment steps, compared with millions of steps required by model-free PPO. Across representation choices, graph structure provides greater robustness under topology-dependent attacks, while simpler representations remain competitive in overall performance. Among successful runs, the number of episodes required to reach the control remains approximately constant as network size increases from 15 to 100 hosts. These results establish learned cyber dynamics as a sample-efficient and scalable basis for autonomous defense, and identify world representation as a central design axis for robustness and scalability.