🤖 AI Summary
This study addresses the challenge of excessive context lengths and high inference overhead in long-horizon agents caused by accumulated reasoning histories, where direct deletion risks altering subsequent interaction trajectories. This work proposes the ICLR method, which reveals a "trajectory amplification effect" when task states are externalized, reframing reasoning as a dynamic working state rather than permanent history to enable safe forgetting of redundant reasoning. Specifically, ICLR achieves training-free, interaction-aware compression by ranking and pruning reasoning blocks online via frozen agent entropy. Evaluated on WorkBuddyBench, the proposed approach increases the average reward to 0.718 while reducing input, output, and cache tokens by 25.5%, 14.4%, and 33.3%, respectively, significantly optimizing long-horizon reasoning efficiency.
📝 Abstract
Long horizon language model agents continually accumulate reasoning history, increasing context length and inference cost even after earlier decisions have been executed and observed. Unlike static Chain of Thought compression, removing historical reasoning can change future actions and the resulting interaction trajectory. We study when such reasoning can be safely forgotten. We propose Interaction Aware Compression for Long Horizon Reasoning (ICLR), a training free online method that ranks reasoning blocks using frozen proxy entropy while preserving actions, tool calls, and observations. On 260 WorkBuddyBench tasks, ICLR improves average reward from 0.699 to 0.718, while reducing input, output, and cache read tokens by 25.5%, 14.4%, and 33.3%, respectively. Ablations reveal trajectory amplification, where local reasoning deletion produces nonlinear changes in total computation by altering subsequent interaction. Representation probing, activation patching, and controlled trajectory analyses further suggest that historical reasoning becomes more replaceable once task relevant derived state has been reliably externalized into code, files, tool outputs, or environmental feedback. These results characterize agent reasoning as dynamic working state rather than permanent interaction history.