🤖 AI Summary
This study investigates whether target alignment scores in the victim space during adversarial attacks can predict target recovery in independent models. To address this, we construct an evidence ladder measurement tool and introduce a supervised referee protocol combined with hybrid control procedures. By integrating robust training techniques such as FARE and TeCoA, we conduct a preregistered study with bootstrap testing to systematically compare six contrastive encoders under matching attacks. Our findings establish that victim-space target scores (VTS) are informative only within fixed robust encoders, prompting a new reporting protocol. Furthermore, we reveal that robust encoders convey more independent evidence than CLIP or SigLIP, although none reach reference-level performance, and no monotonic alignment-evidence relationship exists across encoders.
📝 Abstract
Adversarial attacks on vision-language models optimize an image toward a text target, then cite the attacked model's similarity score as evidence of success. We ask whether that score - victim-space target alignment (VTS) - predicts recovery of the target by an independent model. We first build a measurement instrument: supervised judges outside the attacked geometry, real-target blend controls, shuffled-target negatives, and a reference level derived from a 50% target-image blend. Two preregistered studies then compare six contrastive encoders under a matched attack at three perturbation budgets. Robustly trained encoders (FARE, TeCoA, PMG, TRADES) transfer substantially more independent evidence than vanilla CLIP or SigLIP; all eight contrasts reject at the bootstrap floor. However, no cell reaches the blend-derived reference level. The three best cells fall within its replication band, leaving practical recovery undecided. Within robust encoders, per-sample alignment gain correlates with evidence gain ($ρ= 0.24-0.51$); within vanilla CLIP the correlation is consistent with zero. Across encoders we find no monotone alignment-evidence relation. VTS is therefore informative only within a fixed robust encoder, and we provide a reporting protocol in its place.