TAPDreamer: Transferable Adversarial Patches for World Action Models
This study addresses the vulnerability of visual encoders in world models to localized adversarial attacks and the reliance of existing methods on target outputs. We propose a universal adversarial patch generation method that generalizes across tasks and architectures. By leveraging publicly available encoders, our approach constructs fixed local perturbations through maximizing global representation shifts. It reveals how attention–value interactions enable stable shift broadcasting, facilitating black-box attacks without querying the target policy. Efficient generation is achieved by optimizing the global L1 distance using only six frames from the source task. Evaluated on the LIBERO and RoboTwin benchmarks, the proposed patch reduces task success rates to 0%, underscoring the urgent need to secure shared visual encoders.