๐ค AI Summary
This study addresses the vulnerability of visual encoders in world models to localized adversarial attacks and the reliance of existing methods on target outputs. We propose a universal adversarial patch generation method that generalizes across tasks and architectures. By leveraging publicly available encoders, our approach constructs fixed local perturbations through maximizing global representation shifts. It reveals how attentionโvalue interactions enable stable shift broadcasting, facilitating black-box attacks without querying the target policy. Efficient generation is achieved by optimizing the global L1 distance using only six frames from the source task. Evaluated on the LIBERO and RoboTwin benchmarks, the proposed patch reduces task success rates to 0%, underscoring the urgent need to secure shared visual encoders.
๐ Abstract
World models learn to predict how their environment will evolve, making them an important foundation for general-purpose robotic control. Yet world action models depend on camera inputs whose manipulation can corrupt the visual representations used across tasks and action policies. Existing attacks on these models optimize against the victim's actions or predicted futures and therefore require access to target-model outputs. In this paper, we propose an attack, TAPDreamer, against world action models that instead uses a public encoder alone to construct a fixed local perturbation that transfers across tasks and action architectures. TAPDreamer requires no target-policy queries. Our key insight is that interactions between patch-induced changes in attention weights and value vectors broadcast a nearly identical representation shift far beyond the patch footprint, and this shift remains stable across task observations. Guided by this insight, TAPDreamer uses six frames from one source task to maximize the global L1 distance between clean and patched encoder representations. In closed-loop evaluation, one frozen patch per benchmark, covering about 6.5% of the input, reduces FastWAM's success rate from 97.7% to 0.0% across 40 LIBERO tasks and from 90.8% to 0.0% across 50 RoboTwin tasks; matched random patches retain 81.5% and 79.2% success. The same patches reduce success to 2.1% and 0.8% on two DreamWAM configurations and to 10.0% on Motus. These results show that protecting downstream action generation alone is insufficient: defenses for world action models must also secure shared visual encoders against persistent local perturbations.