🤖 AI Summary
This study addresses the unexplored risk of backdoor vulnerabilities in interactive video generation (IVG) models. We present the first systematic investigation of this threat by proposing BadAction, a novel attack framework that innovatively utilizes action sequences as triggers to implant predefined motion patterns. Under specific trigger conditions, the compromised model outputs frozen frames while disregarding user inputs, yet maintains functional integrity for benign interactions. To further enhance stealthiness, we introduce a multimodal joint poisoning strategy. Experimental results demonstrate that action-only and multimodal triggers achieve average attack success rates of 91.0% and 80.4%, respectively. Moreover, the proposed framework effectively evades existing mainstream backdoor detection defenses, highlighting significant security concerns for deploying IVG systems.
📝 Abstract
Interactive video generation (IVG) models have achieved remarkable progress in producing controllable visual content guided by user-defined actions, yet their security vulnerabilities remain largely unexplored. In this paper, we present the first systematic study of backdoor attacks against the interactivity of IVG models. Based on this attack surface, we propose BadAction, which leverages action-guided triggers to achieve the attack. Specifically, BadAction implants predefined motion patterns into the action sequences of backdoor samples and associates them with a static target video. Once triggered, the backdoored model generates frozen future frames that no longer respond to subsequent user actions, while preserving normal behavior on benign action sequences. In addition, we explore a stealthier attack in which multimodal triggers jointly poison action, text, and image inputs. Experiments show that BadAction achieves average attack success rates of 91.0% with action-only triggers and 80.4% with multimodal triggers. Moreover, extensive defense evaluations show that BadAction successfully bypasses existing backdoor detection methods, revealing a critical security gap in the interactive video generation pipeline. Project page: https://wsad55.github.io/badaction01/.