🤖 AI Summary
This study addresses the high context management costs and frequent omission of critical information encountered by LLM agents during long-horizon tasks. To mitigate these issues, we propose a context optimization method grounded in execution dependency graphs and tool-flow analysis. Specifically, this approach introduces observation reuse signals to integrate execution dependencies into the context rendering mechanism, combining persistent dependency graphs with semantic retrieval for precise information filtering. Experimental results demonstrate that, under a fixed budget of 6K tokens, the proposed method achieves performance comparable to full-history baselines while reducing inference costs by 10.2%–32.2%. These findings indicate that our approach effectively balances information completeness with computational efficiency, offering a practical solution for resource-constrained agentic workflows.
📝 Abstract
LLM agents performing long-horizon tasks accumulate tool results that later steps may need. Passing the full history to every invocation is costly even when it fits within the context window, while reducing it risks omitting needed information. Existing context management methods can overlook how earlier tool results are used in subsequent execution, leaving needed information out of context. We introduce ContextRender, which manages context through a persistent graph of execution dependencies. We develop Tool-Flow Analysis to track how later operations reuse information from earlier tool results, providing a signal called observed reuse. A renderer combines this signal with recency and semantic relevance to select results within a fixed history budget, retaining omitted results for later use. Across AppWorld and 8-objective QA with three execution models, ContextRender outperforms the evaluated context management baselines using a 6K history budget, well below the models' maximum context windows. Within this budget, it achieves task performance close to or above that of passing the full history while reducing mean inference cost by 10.2%-32.2% relative to Full history. Ablations show that observed reuse improves task performance and retention of results reused later.