🤖 AI Summary
This study addresses the problem of task failure in long-running agents caused by the loss of state information. It introduces the concept of Execution Information Requirement (EIR) and proposes the LACUNA framework. Methodologically, this work pioneers an EIR quantitative metric and develops a semantic graph-based real-time measurement approach to decouple operational difficulty from information requirements. LACUNA is employed to isolate and test information dependencies, integrating synthetic task generation with large-scale agent data analysis, while VESTIGE parses information retention and recovery mechanisms within real trajectories. Experimental results demonstrate that the proposed method improves missing information recall accuracy to 100%. Furthermore, it reveals that in failed runs, the rereading rate of relevant information decays significantly as distance increases.
📝 Abstract
Long-running agents must preserve information that later steps depend on. We introduce the Execution Information Requirement (EIR), a lower bound on the information that must remain accessible for correct completion under specified task and access conditions. We develop LACUNA, a framework that generates tasks with known dependencies and varies information demand, retention, and recovery separately from the difficulty of individual operations. Across four models, restoring a missing result raises accuracy on affected recall steps to 100%, compared with 0% for equal-length irrelevant information. Sufficient storage alone does not ensure success: retention policies can discard required results, errors can propagate through later computations, and agents can stop before recovery is complete. We also introduce VESTIGE, which uses agent execution traces to construct semantic graphs and measure information demand for real tasks. Across 72,562 software-agent trajectories, VESTIGE reveals a steeper distance-related decline in solution-relevant rereading for failed runs (RR 0.951 per distance doubling), while adjusted peak demand alone is not associated with failure. Together, these contributions support evaluating whether agents preserve and recover the information their tasks require.