🤖 AI Summary
This study addresses a latent security vulnerability in token compression for large vision models, wherein full-token inference remains robust while compressed pathways are prone to failure. To exploit this, we propose FATA, an adversarial attack integrating attention suppression with cosine similarity-based feature preservation. FATA generates adversarial examples using only visual encoder gradients, eliminating the need to access the deployment-side compressor. This work presents the first conditional blinding attack tailored to compression scenarios. Experiments demonstrate that the proposed method preserves 96.3% accuracy under full-token inference while achieving a 22.1% conditional blinding rate with minimal detectability, effectively balancing attack efficacy and stealthiness. These findings underscore the necessity of joint robustness evaluations encompassing both full and compressed inference pathways.
📝 Abstract
Visual-token compression improves the efficiency of large vision-language models, but can expose failures that full-token evaluation misses. We study adversarial images that preserve full-token correctness yet induce errors after compression, even when both inference paths succeed on the clean image. Creating such failures is challenging because perturbing token importance can also damage the visual content needed for full-token inference. We propose Feature-Aware Token Attack (FATA), which couples attention suppression with cosine-based feature preservation on a fixed set of salient clean-image tokens. In the primary LLaVA-1.5-7B setting, FATA uses only vision-encoder gradients, without access to the deployed compressor, token budget, or downstream task. Across four visually dependent task subsets and four compressors under a controlled reconstruction protocol, FATA achieves SR = 96.3% full-token accuracy retention and CBR = 22.1% conditional blinding, compared with 89.8% and 15.7% for CAA. Ablations support the role of both objectives in balancing compressed-path failure against full-token preservation. FATA also has the lowest measured detection rate among four attacks across three evaluated detectors at a 5% false-positive rate. These findings motivate assessing adversarial robustness jointly across full-token and compressed inference.