🤖 AI Summary
This study addresses goal drift and performance degradation in large language model (LLM) agents caused by context fragmentation during long-horizon tasks. We propose the Adaptive Consistency Graph, a method that incrementally organizes execution evidence within a persistent graph structure and employs an ephemeral view mechanism to construct requirement-centric representations, thereby providing structured and traceable contextual support for decision-making. A key advantage of this approach is its ability to maintain long-term consistency under bounded budgets without replacing the underlying planner. Technically, it integrates graph neural networks with LLM agent architectures. Experimental results demonstrate that our method improves the average success rate of GPT-5.6-luna from 44.5% to 50.2%, yielding particularly significant gains on the BrowseComp-Plus benchmark. The source code has been made publicly available.
📝 Abstract
Large language model agents can often make reasonable local decisions on short tasks, yet their performance degrades when success requires long sequences of dependent actions and tool calls. During execution, task requirements, historical evidence, and the current execution state may gradually become disconnected, so later decisions can drift from the original objective. We study this problem by introducing the Adaptive Consistency Graph (ACG) for long-horizon execution. ACG incrementally organizes execution evidence and its provenance in a persistent graph, then constructs a temporary requirement-centered view for each decision under a bounded context budget. Rather than replacing the base agent's planner or tool executor, ACG provides a structured and traceable context view for each decision. In the matched evaluation, ACG improves GPT-5.6-luna's average success from 44.5\% with ReAct to 50.2\%, with the largest gain on BrowseComp-Plus (73.5\% versus 62.4\%). We further analyze trajectory structure and inference cost to characterize this improvement.