FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the quadratic growth in inference costs and attention dilution in LLM agents caused by historical context accumulation. To this end, it proposes FOCUS, a training-free, decision-preserving context compression framework that reformulates compression as a causal decision preservation problem. By conducting causal analysis over discrete interaction units, FOCUS achieves modular compression dynamically at test time without requiring offline data or fine-tuning. The method is architecture-agnostic and readily applicable as a plug-and-play solution for any closed-source model. Experimental results demonstrate that FOCUS reduces peak context length by 48% and dependency by 73%, while improving task success rates by 8.9 percentage points, thereby establishing a new state of the art.
📝 Abstract
LLM agents accumulate interaction histories that grow linearly with task length, causing quadratic inference cost scaling and performance degradation from attention dilution. Existing context-compression methods learn what to discard offline: by contrastively optimizing guidelines, distilling compressors, or training compression policies. This incurs a substantial cost. Further, the compression policy is learned a priori and is not dynamically conditioned on the evolving test-time trajectories. In this paper we ask a complementary question: Which past interactions causally shape the agent's future decisions? We recast context compression as a causal decision preservation problem over discrete interaction units and introduce FOCUS, a training-free context compression framework that operates entirely at test time. Our method requires no offline data collection or fine-tuning, and is architecture-agnostic, attaching to any closed-API frontier model as a modular compression layer. We evaluate FOCUS on diverse agentic benchmarks including API and tool-calling, QA, web domain and multi-turn dialogue. Our method establishes new state of the art performance, cutting peak context by up to 48% and dependency by 73% while improving task success by up to 8.9 percentage points over uncompressed execution.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
context compression
inference cost
attention dilution
decision preservation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-Free Context Compression
Causal Decision Preservation
LLM Agents
Test-Time Compression
Architecture-Agnostic
🔎 Similar Papers