Understanding and Mitigating Token-Pruning-Induced Vulnerabilities in VLMs

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the security vulnerabilities introduced by token pruning in vision-language models (VLMs), which, despite accelerating inference, induces attention collapse and amplifies malicious semantics. We present the first systematic evaluation of pruning safety, revealing that query compression unexpectedly enhances security while elucidating the underlying mechanism of malicious amplification. Building on these insights, this work proposes Security-Aware Pruning (SAP), a plug-and-play method that ensures runtime protection through malicious anchor identification, benign token restoration, and attention redistribution. Experimental results demonstrate that SAP reduces the attack success rate by 62%, effectively mitigating the security deficiencies of token pruning without compromising inference efficiency or general-purpose performance.
📝 Abstract
Token-Pruning accelerates Vision-Language Models by removing redundant visual tokens, yet its safety implications remain underexplored. In this work, we present the first comprehensive safety evaluation of Token-Pruning mechanisms and find that: most pruning strategies significantly degrade safety as pruning ratios increase, whereas Query-based Compression shows the opposite, with extreme pruning (up to 99.8%), unexpectedly improves model safety. This sharp contrast prompts a key question: How do different Token-Pruning strategies reshape model safety behavior, and is it possible to enhance safety without sacrificing acceleration? To answer this, we identify an unrecognized mechanism, termed Pruning-Induced Malicious Amplification, where removal of background tokens triggers a side effect: forcing the model's attention to collapse onto a few retained malicious anchors within the foreground, inadvertently amplifying their toxic semantics under jailbreak. To address that, we propose an inference-time and plug-and-play Safety-Aware Pruning (SAP) mechanism that counteracts such dominance via three steps: (1) identifying malicious anchors, (2) restoring pruned benign tokens, and (3) reallocating excessive attention from malicious anchors to benign tokens. Extensive experiments across three safety and four utility benchmarks demonstrate that SAP mitigates pruning-induced vulnerabilities, i.e., reducing ASR by up to 62%, without compromising efficiency or utility.
Problem

Research questions and friction points this paper is trying to address.

Token Pruning
Vision-Language Models
Safety Vulnerabilities
Jailbreak
Malicious Amplification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Token Pruning
Vision-Language Models
Safety Evaluation
Pruning-Induced Malicious Amplification
Safety-Aware Pruning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shuailong Wang
University of Electronic Science and Technology of China, Chengdu, China
Xinyu Lyu
Xinyu Lyu
Southwestern University of Finance and Economics
AI safetyMultimedia Learning
S
Shengming Yuan
University of Electronic Science and Technology of China, Chengdu, China
J
Jingkuan Song
Tongji University, Shanghai, China; Shanghai Innovation Institute, Shanghai, China
H
Heng Tao Shen
Tongji University, Shanghai, China
Lianli Gao
Lianli Gao
UESTC
Vision and Language