🤖 AI Summary
This study addresses the loss of task requirements and file states caused by limited context windows when large language model (LLM) agents interact directly with workspaces. To mitigate this, we propose RunningTab, a framework that introduces an environment-side tab-based persistent recording mechanism to decouple agent memory from environmental states. This approach dynamically tracks task progress, file read/write statuses, and unprocessed candidates in real time, enabling continuous alignment between task requirements and workspace content. Extensive experiments conducted across three benchmarks and three LLMs demonstrate that RunningTab consistently outperforms both direct interaction paradigms and existing baselines, significantly enhancing the completeness of delivered artifacts.
📝 Abstract
Much knowledge work produces new deliverables from files a workspace already holds, and LLM agents are beginning to take such work over. Through direct corpus interaction, an agent can search and read any of those files from a terminal with no indexing, and producing a deliverable from many of them in this way is what we call direct workspace interaction (DWI). Reaching the files, however, is only half the task: nothing keeps track of what the task asks for, what has been read, and what was listed but never opened, all of which slip through the context window without leaving a trace, so an agent may extract a figure and still deliver a report without it. To address this, we present RunningTab, a framework that equips direct workspace interaction with an environment-side tab: a per-task record of what the task still owes, kept by the environment alongside the agent. Specifically, the agent adds its requirements, while the environment records every file read as an excerpt with its provenance and every listed but unopened file as a candidate; the agent can then see each requirement beside its best-matching excerpts and top unopened candidates, resolve it against matching content or set it aside with a reason, and, should it try to finish with requirements still open, receive them in a finish check. We validate RunningTab on three benchmarks with three LLMs, where it consistently outperforms plain DWI and baselines that keep the record in the model, while its tab usually holds the values a deliverable needs once seen.