🤖 AI Summary
This study addresses the limitation of existing video jailbreak attacks that treat videos merely as carriers while overlooking scene context as a critical attack surface. We present the first systematic investigation of this vulnerability, proposing a scene-aware adaptive black-box attack framework. By integrating adaptive scene construction, scene-aware prompt search, and black-box feedback optimization, our method leverages video contexts to dynamically induce multimodal large language models into generating harmful content. Experimental results demonstrate that the proposed framework achieves an average attack success rate of 91.5%, outperforming baselines by 29.1%. Notably, it maintains a 72.3% success rate even under stringent defense mechanisms, significantly revealing the latent security risks associated with video scene contexts in multimodal systems.
📝 Abstract
Video Multimodal Large Language Models (Video-MLLMs) support reasoning over video inputs, yet remain vulnerable to jailbreak attacks that elicit policy-violating responses. Existing video jailbreaks primarily manipulate how harmful queries are visually presented, thereby treating video merely as a carrier. Consequently, the surrounding video scenario remains unexplored as a contextual attack surface. In this paper, we show that the same harmful query can elicit different safety responses when placed in different video scenarios.
To systematically exploit this vulnerability, we propose SceneJail, an adaptive black-box jailbreak framework with two coordinated components. Adaptive Scenario Construction dynamically searches for a surrounding scenario that is contextually compatible with the harmful query. Scenario-aware Prompt Search uses black-box response feedback to search for textual guidance tailored to the selected scenario. Extensive evaluations on the HADES and SafeBench datasets across eight Video-MLLMs, including two proprietary models, GPT-4.1 and Gemini3.5-Flash, demonstrate the effectiveness of SceneJail. SceneJail-F, which presents the complete query persistently, achieves average attack success rates (ASR) up to 91.5%, outperforming the strongest baselines by 29.1 percentage points. Furthermore, SceneJail-S, which distributes the query across successive frames, remains highly robust against current defenses, retaining a 72.3% ASR even under strict image filtering.