TooBadRL: Trigger Optimization to Boost Effectiveness of Backdoor Attacks on Deep Reinforcement Learning

📅 2025-06-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing DRL backdoor attacks rely on manually designed, simplistic triggers, neglecting joint spatiotemporal-amplitude optimization—leading to low attack efficacy and poor stealth. This paper proposes the first three-axis co-optimization framework for DRL backdoor triggers: (1) a performance-aware adaptive injection timing freezing mechanism for precise key-frame triggering; (2) a Shapley-value-based cooperative game-theoretic dimension selection method to dynamically identify high-impact state dimensions; and (3) an environment-constrained, gradient-driven amplitude optimization strategy to balance attack potency and behavioral naturalness. Evaluated across three mainstream RL algorithms—PPO, SAC, and A2C—on nine standard benchmark tasks, our approach achieves an average 37.2% improvement in attack success rate while degrading clean-task performance by less than 0.8%, significantly outperforming state-of-the-art methods.

Technology Category

Search and Optimization: Learning to SearchIntelligent Robots: Learning & Optimization for ROBMultiagent Systems: Adversarial Agents

Application Category

User Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsResponsible Web: Machine-in-the-loop, human agency and autonomySearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
Deep reinforcement learning (DRL) has achieved remarkable success in a wide range of sequential decision-making domains, including robotics, healthcare, smart grids, and finance. Recent research demonstrates that attackers can efficiently exploit system vulnerabilities during the training phase to execute backdoor attacks, producing malicious actions when specific trigger patterns are present in the state observations. However, most existing backdoor attacks rely primarily on simplistic and heuristic trigger configurations, overlooking the potential efficacy of trigger optimization. To address this gap, we introduce TooBadRL (Trigger Optimization to Boost Effectiveness of Backdoor Attacks on DRL), the first framework to systematically optimize DRL backdoor triggers along three critical axes, i.e., temporal, spatial, and magnitude. Specifically, we first introduce a performance-aware adaptive freezing mechanism for injection timing. Then, we formulate dimension selection as a cooperative game, utilizing Shapley value analysis to identify the most influential state variable for the injection dimension. Furthermore, we propose a gradient-based adversarial procedure to optimize the injection magnitude under environment constraints. Evaluations on three mainstream DRL algorithms and nine benchmark tasks show that TooBadRL significantly improves attack success rates, while ensuring minimal degradation of normal task performance. These results highlight the previously underappreciated importance of principled trigger optimization in DRL backdoor attacks. The source code of TooBadRL can be found at https://github.com/S3IC-Lab/TooBadRL.
Problem

Research questions and friction points this paper is trying to address.

Optimizing backdoor triggers in DRL attacks
Enhancing attack success via temporal, spatial, magnitude optimization
Addressing simplistic trigger flaws in DRL backdoor strategies
Innovation

Methods, ideas, or system contributions that make the work stand out.

Performance-aware adaptive freezing for timing
Shapley value for influential dimension selection
Gradient-based adversarial magnitude optimization
🔎 Similar Papers
No similar papers found.