StateTape: Action-Conditioned Evidence Lifecycle Modeling for Long-Horizon Coding Agents

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the failure of long-horizon coding agents caused by unbounded context growth and the inability of conventional methods to perceive evolving code dependencies. We propose StateTape, a framework that models code repositories as symbol-level code graphs and dynamically rewrites agent contexts by marking written changes. To our knowledge, this is the first approach to identify stale records based on code structure rather than textual inference, enabling action-conditioned evidence lifecycle management. We further introduce TraceBench, a companion benchmark for evaluation. Experiments across six agents and three benchmarks demonstrate that StateTape significantly improves resolution rates by effectively purging spurious records and precisely retrieving critical information, all while incurring minimal computational overhead.
📝 Abstract
Despite the recent success of coding agents built on large language models, it remains challenging to run them over long horizons, since every observation is appended to the context and the context grows with each one. History-based maintenance is a common remedy, which masks or summarizes old observations, or prunes what a model reads as useless, and bounds the context at little cost. However, it decides from the text of the history alone and sees nothing of how the code is connected. Since a coding agent edits code many times over a single task, and each write can change what code elsewhere means, such maintenance may keep records a write has falsified, drop ones that still hold, and miss code the agent needs next. To overcome these challenges, this paper proposes StateTape, a novel and scalable framework that rewrites a coding agent's context as the repository changes rather than as the context grows. The key idea of StateTape is to model the repository as a symbol-level code graph, whose dependencies and language rules expose which symbols a write can affect. Upon this graph, a tape marks the symbols each write changed, which turns staleness from an inference about text into an observation of the agent's writes. We propose a per-write procedure in which the tape nominates the records a write could have falsified while a small manager model settles what the write log cannot, and further provide a theoretical analysis and TraceBench, a benchmark that labels what an agent is holding against what is actually needed. Empirically, we demonstrate that StateTape can effectively clear falsified records and retrieve what is needed, and thus achieve a higher resolve rate in all experiments spanned by six coding agents and three edit-heavy benchmarks with little computational overhead.
Problem

Research questions and friction points this paper is trying to address.

coding agents
long-horizon tasks
context management
history maintenance
code dependencies
Innovation

Methods, ideas, or system contributions that make the work stand out.

coding agents
context management
symbol-level code graph
evidence lifecycle modeling
long-horizon tasks