🤖 AI Summary
This study addresses the underexplored risk that visual token compression in large vision-language models introduces specific adversarial failure modes. To this end, it formally defines “compression-induced failures” and proposes CIRA, a novel attack framework. CIRA pioneers an attribution analysis mechanism to disentangle inherited vulnerabilities from compression-specific risks, leveraging white-box optimization of encoder objectives to manipulate token priorities for targeted attacks without downstream labels or white-box access during deployment. Experimental results demonstrate that CIRA achieves an average attack success rate of 20.35%, substantially outperforming the full-token baseline of 6.92%. These findings confirm the independence of compression-induced risks and expose critical limitations in existing defense strategies such as cross-view selection, highlighting urgent security challenges in efficient vision-language modeling.
📝 Abstract
Visual token compression reduces the inference cost of Large Vision-Language Models (LVLMs). However, aggregate robustness measures do not reveal whether a particular adversarial failure is induced by compression or inherited from the underlying model. We define a compression-specific failure (CSF) as an adversarial input that remains correct under full-token inference but fails after compression, casting compression-induced risk as a paired failure attribution problem. Within a controlled diagnostic cohort, counterfactuals show that retained-set allocation causally changes compressed correctness and reveal a negative association between recovery and representation drift in displaced evidence. Motivated by these findings, we propose CIRA, a Compression-Induced Risk Attack for Large Vision-Language Models. Under a vision-encoder white-box setting, CIRA optimizes image perturbations through encoder-side objectives that manipulate token priorities across candidate compression budgets while preserving displaced evidence. CIRA uses no downstream questions or labels and requires no access to the language model, deployed compressor, or exact compression budget. Across 12 dataset-compressor settings evaluated at four budgets, CIRA achieves a mean CSFR of 20.35% while limiting full-token attack success to 6.92%, with similar behavior on additional LVLM families. A cross-view selection-stabilization defense substantially suppresses CIRA, although Adaptive CIRA partially restores its effectiveness. These results show that compression-specific failures persist under restricted access and support paired evaluation of full-token and compressed inference for attributing risk to visual-token compression.