BadAction: Backdoor Attacks on Interactive Video Generation via Action-Guided Triggers

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unexplored risk of backdoor vulnerabilities in interactive video generation (IVG) models. We present the first systematic investigation of this threat by proposing BadAction, a novel attack framework that innovatively utilizes action sequences as triggers to implant predefined motion patterns. Under specific trigger conditions, the compromised model outputs frozen frames while disregarding user inputs, yet maintains functional integrity for benign interactions. To further enhance stealthiness, we introduce a multimodal joint poisoning strategy. Experimental results demonstrate that action-only and multimodal triggers achieve average attack success rates of 91.0% and 80.4%, respectively. Moreover, the proposed framework effectively evades existing mainstream backdoor detection defenses, highlighting significant security concerns for deploying IVG systems.
📝 Abstract
Interactive video generation (IVG) models have achieved remarkable progress in producing controllable visual content guided by user-defined actions, yet their security vulnerabilities remain largely unexplored. In this paper, we present the first systematic study of backdoor attacks against the interactivity of IVG models. Based on this attack surface, we propose BadAction, which leverages action-guided triggers to achieve the attack. Specifically, BadAction implants predefined motion patterns into the action sequences of backdoor samples and associates them with a static target video. Once triggered, the backdoored model generates frozen future frames that no longer respond to subsequent user actions, while preserving normal behavior on benign action sequences. In addition, we explore a stealthier attack in which multimodal triggers jointly poison action, text, and image inputs. Experiments show that BadAction achieves average attack success rates of 91.0% with action-only triggers and 80.4% with multimodal triggers. Moreover, extensive defense evaluations show that BadAction successfully bypasses existing backdoor detection methods, revealing a critical security gap in the interactive video generation pipeline. Project page: https://wsad55.github.io/badaction01/.
Problem

Research questions and friction points this paper is trying to address.

Interactive Video Generation
Backdoor Attacks
Security Vulnerabilities
Action-Guided Triggers
Multimodal Triggers
Innovation

Methods, ideas, or system contributions that make the work stand out.

Backdoor Attack
Interactive Video Generation
Action-Guided Triggers
Multimodal Poisoning
Adversarial Robustness
Z
Zhihang Wu
Key Laboratory of AI Safety of CAS, Institute of Computing Technology, Chinese Academy of Sciences (CAS), Beijing, China; University of Chinese Academy of Sciences, Beijing, China
Zhongqi Wang
Zhongqi Wang
Institute of Computing Technology, Chinese Academy of Sciences
Model Robustness
J
Jie Zhang
Key Laboratory of AI Safety of CAS, Institute of Computing Technology, Chinese Academy of Sciences (CAS), Beijing, China; University of Chinese Academy of Sciences, Beijing, China
F
Fengming Gu
Key Laboratory of AI Safety of CAS, Institute of Computing Technology, Chinese Academy of Sciences (CAS), Beijing, China; University of Chinese Academy of Sciences, Beijing, China; School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences, Beijing, China
Shiguang Shan
Shiguang Shan
Professor of Institute of Computing Technology, Chinese Academy of Sciences
Computer VisionPattern RecognitionMachine LearningFace Recognition
Xilin Chen
Xilin Chen
Institute of Computing Technology, Chinese Academy of Sciences
Computer VisionPattern RecognitionMachine Learning