Crossing the Cyber Divide: Sim-to-Sim and Sim-to-Real Transfer for RL Agents

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of transferring cyber-attack agent policies across simulators and real-world environments by proposing a unified framework that decouples state alignment from action translation. The proposed method formulates both simulation-to-simulation and simulation-to-real policy transfer as spatial alignment problems, integrating reinforcement learning with state-action mapping mechanisms to achieve zero-shot policy transfer without retraining. Experimental evaluations across four platforms validate the feasibility of this framework: transferred agents either preserve their source-domain performance or attain a 45.2% win rate while exhibiting exceptionally high behavioral similarity. These results demonstrate that the approach effectively overcomes the bottlenecks associated with cross-platform deployment.
📝 Abstract
Cyber attack agents are typically trained and evaluated within a single simulator, making it unclear whether learned policies transfer beyond the environments in which they were developed. This limitation hinders both deployment and fair comparison, as cyber simulators differ substantially in their state representations, observation models, and action spaces. In this paper, we study policy transfer across cyber environments and argue that simulator-to-simulator and simulator-to-real transfer can be viewed as instances of the same underlying alignment problem. We propose a framework that separates state alignment from action translation, enabling a policy trained in one environment to operate in another without retraining. We evaluate transfer across four cyber platforms, CyberBattleSim, NetSecGame, CyberWheel, and NASim, including emulated deployments in NASim. Our experiments show that zero-shot transfer is feasible, fully preserving source-policy performance in closely aligned environments and achieving 45.2% win rates when transferring policies whose source performance is 60.5%. In emulated virtual machine environments, transferred policies exhibit a Jensen-Shannon divergence of 0.085 from native policies, indicating strong behavioral similarity. Code and benchmarks are available at: https://anonymous.4open.science/r/RL-Transfer-between-env-4F47/.
Problem

Research questions and friction points this paper is trying to address.

Cyber attack agents
Sim-to-Real transfer
Sim-to-Sim transfer
Reinforcement learning
Policy transfer
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sim-to-Real Transfer
Zero-Shot Transfer
Reinforcement Learning
State Alignment
Cyber Security
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.