PreAct: Computer-Using Agents that Get Faster on Repeated Tasks

📅 2026-06-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing computer operation agents inefficiently re-reason from scratch even when repeatedly executing the same tasks. This work proposes a reusable program memory mechanism that compiles the first successful execution of a task into a finite-state-machine program, which is then directly replayed in subsequent runs, relinquishing control back to the agent only upon detecting anomalous screen states. The approach integrates state-machine compilation, screen-state matching, an independent evaluator for correctness verification, and a program selection strategy combining language models with embedding-based retrieval, ensuring reliability and enabling graceful fallback upon failure. Evaluated across mobile, desktop, and web benchmarks, the system achieves 8.5–13× speedup and completes 1.75–2.6 additional tasks on average per benchmark, substantially outperforming strong baseline replay methods.
📝 Abstract
Computer-using agents drive real software through the screen -- clicking and typing -- but they solve every task from scratch: asked to repeat a task, an agent re-reads the screen, re-reasons every tap, and pays the full cost again. We present PreAct, which lets such an agent get faster on tasks it has done before. The first time it succeeds, PreAct compiles the run into a small state-machine program-states that check the screen, transitions that act-and on later runs replays it directly instead of invoking the agent 8.5-13x faster, with no per-step language-model calls. Replay is not blind: at each step PreAct checks that the screen matches what the program expects before acting, and hands control back to the agent the moment something is off. PreAct applies the same discipline when deciding what to keep: a freshly compiled program enters the store only if, re-run from a clean state, an independent evaluator confirms it solved the task-catching programs that replay to their last step yet leave the task undone. Across a mobile, a desktop, and a web benchmark, this store-time check separates repeated runs that improve from ones that degrade as faulty programs accumulate, worth 1.75-2.6 tasks per benchmark, the same direction on all three; a fallback that explores afresh when no program fits brings PreAct level with a strong record-and-replay baseline. We also report what did not matter: prompt wording, runtime guardrails, and whether a language model or a plain embedding retriever selects which program to reuse.
Problem

Research questions and friction points this paper is trying to address.

computer-using agents
task repetition
execution efficiency
screen-based interaction
agent acceleration
Innovation

Methods, ideas, or system contributions that make the work stand out.

PreAct
state-machine compilation
replay acceleration
screen-based agent
program validation
🔎 Similar Papers
No similar papers found.