PastForward: Faster On-Device GUI Agents via Computational Experience Reuse

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high per-action inference latency of vision-language models when deploying GUI agents on edge devices. To mitigate this, we propose a fine-grained computational experience reuse mechanism that requires no fine-tuning. During decoding, historical sequences are retrieved as multi-token proposals and verified in a single forward pass. Across steps, GUI state transitions are leveraged to initiate speculative inference early; prior computations are retained and KV caches propagated only when predictions match observations. Evaluated on the AndroidWorld benchmark, our approach achieves a 1.63–2.36× inference speedup without compromising task success rates, effectively facilitating device-adaptive deployment in dynamic environments.
📝 Abstract
Running GUI agents on edge devices can keep sensitive screens and interaction histories local, but the computational cost of inference at every action step makes deployment challenging. Existing GUI agent systems either perform full vision-language model (VLM) inference at each action step or reuse coarse-grained knowledge matched to prior tasks. However, dynamic mobile environments and user tasks make it difficult to fully utilize prior task executions without additional fine-tuning or task-specific offline exploration. To address this challenge, we present PastForward, a system that accelerates GUI agents through validated, fine-grained reuse of computational experience accumulated during ordinary task execution. During decoding, PastForward retrieves prior output sequences as device-adaptive multi-token proposals and verifies them in a single VLM forward pass. Across action steps, it uses prior GUI transitions to begin next-step inference while the device executes the current action, retains the early computation only when the predicted screen matches the observed screen, and carries reusable KV states forward. We evaluate PastForward on AndroidWorld workloads derived from real mobile usage patterns using multiple VLM backbones across server and edge platforms. On device, PastForward achieves action-step latency speedups of 1.63-2.36$\times$ while maintaining task success rates.
Problem

Research questions and friction points this paper is trying to address.

GUI agents
edge devices
computational cost
experience reuse
inference acceleration
Innovation

Methods, ideas, or system contributions that make the work stand out.

GUI Agents
Computational Experience Reuse
On-Device Inference
Vision-Language Model
KV State Reuse