🤖 AI Summary
This study addresses the unpredictable robustness of CLIP under mask pruning by identifying "spurious inversion" as the critical factor underlying unstable masking performance. To this end, we introduce the Spurious Inversion Metric (SIM) and propose a label-free, pre-deployment diagnostic framework that integrates semantic masking, text similarity analysis, and asynchronous GPU batch partitioning for efficient evaluation. Experimental results demonstrate that SIM significantly predicts masking efficacy, enabling optimized models to match or surpass baseline performance. This work thereby offers a reliable solution for the robust deployment of compressed CLIP architectures.
📝 Abstract
This paper demonstrate that whether masking-based token pruning helps or hurts worst-group robustness can be predicted before deployment, without labels or fine-tuning. A systematic study of semantic masking across 8 spurious-correlation benchmarks shows its effect on worst-group accuracy is highly unstable: it improves accuracy by up to 82.5\% relative on some datasets and degrades it by up to 100\% on others. We trace this instability to spurious inversion: background patches receive higher CLIP text-similarity than the true object when the spurious attribute is background-separable, inverting the assumption every text- and attention-guided pruning method relies on. We introduce the Spurious Inversion Metric (SIM), a label-free, pre-deployment diagnostic whose sign predicts this effect with statistical significance (binomial $p=0.035$) across all 8 datasets, and remains dependable across 6 CLIP architectures with a clean foreground/background split. Naive masking is itself a major source of risk: it causes the largest average-accuracy loss of any method we evaluate, and its own per-image segmentation step is a significant runtime bottleneck. To address this, we design a batched, synchronization-free GPU segmentation routine that cuts this overhead from 3.5$\times$ to 1.75$\times$ baseline. Gating deployment by SIM's sign recovers masking's benefits while avoiding its worst failures, matching or exceeding a strong pruning baseline on 7 of 8 datasets.