🤖 AI Summary
This study addresses the high system latency caused by idle waiting in LLM agents during long-duration tool invocations. To mitigate this, it pioneers the adaptation of processor out-of-order execution mechanisms to LLM agents by proposing an optimization framework based on out-of-order speculative execution. This framework proactively simulates future actions in parallel within copy-on-write sandboxes, employing dependency tracking and committed-state verification techniques to ensure strict state consistency. Evaluated on benchmarks including SWE-bench, the proposed approach achieves 1.31× to 1.35× speedups with zero false acceptances across 4,010 verifications, effectively enhancing both the operational efficiency and reliability of LLM agents.
📝 Abstract
Long-running tools can dominate coding-agent latency: compilers, test suites, and repository commands take seconds to minutes while the agent idles. This observation stall presents the same tension that drove out-of-order processors -- asequential interface hides work that can be predicted and started early, but a speculative result may become visible only after it and every earlier step have been validated.
We present TomasuLLM, a runtime that executes agent tool calls out of trajectory order while preserving task-execution correctness. It drafts future actions, runs them in isolated copy-on-write sandboxes, traces their dependencies and effects, and commits results in trajectory order only after validation against committed state. Across three benchmarks spanning sub-second to minutes-long tool calls, TomasuLLM improves the reported benchmark means and scales with tool latency: 1.31x on 100 SWE-bench Verified tasks, 1.35x on 28 Terminal-Bench 2.0 tasks, and 1.27x matched progress on 18 SWE-Marathon sessions. Across 4,010 audited commit-validation records, it produces zero false accepts.