Boosting Transferable Adversarial Attacks against Deep Reinforcement Learning

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in deep reinforcement learning where black-box transfer attacks often underperform random noise perturbations. To overcome this limitation, we propose a trajectory-level attack method based on differentiable environment models. By transcending single-step optimization constraints, our approach employs a temperature smoothing strategy to approximate non-differentiable agent policies and integrates receding-horizon optimization with FGSM to generate globally optimal perturbation sequences. Experimental evaluations on the CartPole-v1 environment demonstrate that the proposed method significantly outperforms existing baselines across white-box, cross-model, and cross-algorithm settings, achieving highly transferable attacks with strong generalization capabilities.
📝 Abstract
Most adversarial attacks on deep reinforcement learning (DRL) assume white-box access to the victim policy, which rarely holds in practice. This paper studies transfer-based black-box attacks on DRL: the attacker crafts observation perturbations on a white-box surrogate agent and feeds them to an unknown victim. We formulate the attack as return minimization under a per-step perturbation budget. We first show that transplanting transferable image-classification attacks (FGSM, MI-FGSM, and NI-FGSM) with a per-step objective yields perturbations that transfer but are no stronger than random noise of the same budget. We then propose a trajectory-level attack that optimizes a sequence of perturbations over a receding horizon through a differentiable model of the environment and a temperature-smoothed surrogate policy, with the same optimizers. On CartPole-v1 with ten DQN and DDQN agents and 100 surrogate--victim pairs, the trajectory-level attack outperforms per-step attacks and random noise in the white-box, cross-model, and cross-algorithm settings.
Problem

Research questions and friction points this paper is trying to address.

Adversarial Attacks
Deep Reinforcement Learning
Transferability
Black-box Attack
Innovation

Methods, ideas, or system contributions that make the work stand out.

Transferable Adversarial Attacks
Deep Reinforcement Learning
Trajectory-level Attack
Black-box Attack
Differentiable Environment Model
🔎 Similar Papers
No similar papers found.