🤖 AI Summary
This study addresses the security vulnerability wherein GUI agents struggle to identify malicious state transitions relying solely on screen observations, and proposes the LGWM model. This work pioneers a decoder-free, action-conditioned representation space prediction paradigm that reformulates safety verification as vector similarity comparison. Trained on 1.85 million real-world GUI transitions without semantic annotations, LGWM efficiently detects anomalies and establishes world models in a novel role as safety verifiers. On the RSWT-BENCH benchmark, it achieves an AUC of 0.987 with merely 17 milliseconds of decision latency. This performance rivals that of top-tier vision-language models while being three orders of magnitude faster, and effectively distinguishes harmful violations from benign interactions.
📝 Abstract
A login screen that appears after a tap on Sign in is expected; the same screen after a tap on View order is an attack. For GUI agents, safety is therefore a property of the transition rather than of the screen, and a monitor that inspects only screens can be defeated by reusing a legitimate one. Judging a transition requires an expectation of what should have followed the action. Existing GUI world models provide one, but they output it as text, code, or images, so checking it against the observed screen requires a second model to judge the two. We argue that a world model meant for verification should instead predict in the space in which observations are encoded, and present LGWM, a decoder-free, action-conditioned world model that predicts the representation of the next screen directly, trained without semantic annotation on 1.85M real GUI transitions. Verification reduces to a vector comparison, and the same signal reveals whether a mismatch is harmful. We evaluate on RSWT-BENCH, a diagnostic where each credential screen appears under both a legitimate and a hijacked transition, so detectors that see only the screen are at chance by construction. The training-free score reaches 0.987 AUC at 17 ms per decision, on par with the strongest closed-source VLMs and about ten AUC points above generative GUI world models at over three orders of magnitude lower latency. The residual direction reaches 0.953 AUC at separating harmful from benign violations, where prompted VLMs are near chance. Further analyses show that the prediction is a usable future state rather than an anomaly score. World models have mostly served as simulators or planners; our results point to a third role, verification, for which predicting in representation space is the natural design.