Towards Generalizable Autonomous Penetration Testing via Domain Randomization and Meta-Reinforcement Learning

๐Ÿ“… 2024-12-05
๐Ÿ›๏ธ arXiv.org
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the poor generalization and limited transferability of reinforcement learning (RL)-driven automated penetration testing in real-world scenarios. To this end, we propose the GAP framework, establishing a Real-to-Sim-to-Real closed-loop pipeline that enables zero-shot cross-environment transfer and minute-scale adaptation. GAP pioneers the integration of domain randomization into autonomous penetration testing, leveraging large language models to generate diverse, semantically plausible synthetic vulnerability environments. Furthermore, it synergistically combines domain randomization with meta-reinforcement learning to enhance policy robustness and generalization against unseen target systems. Experimental results demonstrate that GAP significantly mitigates sim-to-real performance degradation across multiple real-world vulnerable virtual machines. It supports policy learning in unknown environments, zero-shot transfer to semantically similar environments, and rapid adaptation to heterogeneous targetsโ€”achieving practical deployability for adaptive red-teaming.

Technology Category

Machine Learning: Reinforcement LearningNatural Language Processing: Safety and RobustnessComputer Vision: Adversarial Attacks & Robustness

Application Category

Responsible Web: Machine-in-the-loop, human agency and autonomySearch and Retrieval-Augmented AI: Agentic searchUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systems
๐Ÿ“ Abstract
With increasing numbers of vulnerabilities exposed on the internet, autonomous penetration testing (pentesting) has emerged as an emerging research area, while reinforcement learning (RL) is a natural fit for studying autonomous pentesting. Previous research in RL-based autonomous pentesting mainly focused on enhancing agents' learning efficacy within abstract simulated training environments. They overlooked the applicability and generalization requirements of deploying agents' policies in real-world environments that differ substantially from their training settings. In contrast, for the first time, we shift focus to the pentesting agents' ability to generalize across unseen real environments. For this purpose, we propose a Generalizable Autonomous Pentesting framework (namely GAP) for training agents capable of drawing inferences from one to another -- a key requirement for the broad application of autonomous pentesting and a hallmark of human intelligence. GAP introduces a Real-to-Sim-to-Real pipeline with two key methods: domain randomization and meta-RL learning. Specifically, we are among the first to apply domain randomization in autonomous pentesting and propose a large language model-powered domain randomization method for synthetic environment generation. We further apply meta-RL to improve the agents' generalization ability in unseen environments by leveraging the synthetic environments. The combination of these two methods can effectively bridge the generalization gap and improve policy adaptation performance. Experiments are conducted on various vulnerable virtual machines, with results showing that GAP can (a) enable policy learning in unknown real environments, (b) achieve zero-shot policy transfer in similar environments, and (c) realize rapid policy adaptation in dissimilar environments.
Problem

Research questions and friction points this paper is trying to address.

Enhance generalization in autonomous penetration testing.
Bridge gap between simulated and real environments.
Improve policy adaptation in unseen scenarios.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Domain randomization for realistic simulation
Meta-reinforcement learning for generalization
Real-to-Sim-to-Real pipeline integration
๐Ÿ”Ž Similar Papers
No similar papers found.
National University of Defense Technology | Tsinghua University
Shicheng Zhou
Shicheng Zhou
Unknown affiliation
J
Jingju Liu
College of Electronic Engineering, National University of Defense Technology, Hefei 230037, China; Institute for Network Sciences and Cyberspace, Tsinghua University, Beijing 100084, China; Beijing National Research Center for Information Science and Technology, Tsinghua University, Beijing 100084, China
Y
Yuliang Lu
College of Electronic Engineering, National University of Defense Technology, Hefei 230037, China
Jiahai Yang
Jiahai Yang
Tsinghua University
Y
Yue Zhang
College of Computer Science and Technology, National University of Defense Technology, Changsha 410073, China
J
Jie Chen
College of Electronic Engineering, National University of Defense Technology, Hefei 230037, China