🤖 AI Summary
This study addresses the quadratic growth in inference costs and attention dilution in LLM agents caused by historical context accumulation. To this end, it proposes FOCUS, a training-free, decision-preserving context compression framework that reformulates compression as a causal decision preservation problem. By conducting causal analysis over discrete interaction units, FOCUS achieves modular compression dynamically at test time without requiring offline data or fine-tuning. The method is architecture-agnostic and readily applicable as a plug-and-play solution for any closed-source model. Experimental results demonstrate that FOCUS reduces peak context length by 48% and dependency by 73%, while improving task success rates by 8.9 percentage points, thereby establishing a new state of the art.
📝 Abstract
LLM agents accumulate interaction histories that grow linearly with task length, causing quadratic inference cost scaling and performance degradation from attention dilution. Existing context-compression methods learn what to discard offline: by contrastively optimizing guidelines, distilling compressors, or training compression policies. This incurs a substantial cost. Further, the compression policy is learned a priori and is not dynamically conditioned on the evolving test-time trajectories. In this paper we ask a complementary question: Which past interactions causally shape the agent's future decisions? We recast context compression as a causal decision preservation problem over discrete interaction units and introduce FOCUS, a training-free context compression framework that operates entirely at test time. Our method requires no offline data collection or fine-tuning, and is architecture-agnostic, attaching to any closed-API frontier model as a modular compression layer. We evaluate FOCUS on diverse agentic benchmarks including API and tool-calling, QA, web domain and multi-turn dialogue. Our method establishes new state of the art performance, cutting peak context by up to 48% and dependency by 73% while improving task success by up to 8.9 percentage points over uncompressed execution.