🤖 AI Summary
This study addresses the evaluation ambiguity arising from coupled decisions in context compression strategies for LLM agents by systematically decoupling, for the first time, three decision dimensions: compression timing, content, and magnitude. Through 35,000 experiments conducted across multiple open-source models on the SWE-bench and Terminal-Bench benchmarks, this work quantifies the impact of each dimension on task success rate, latency, and cost. The results reveal that token reduction does not necessarily yield efficiency or cost benefits; aggressive compression can even increase latency by up to 80%. Furthermore, compression effectiveness is highly contingent upon model and task characteristics. These findings underscore the necessity of tailoring compression strategies to specific scenarios to effectively balance performance and computational cost.
📝 Abstract
As LLM agents tackle longer tasks, they increasingly compress growing histories of reasoning, actions, and tool outputs. Compression can reduce token use, but it also changes the information available for later decisions. Existing agentic harnesses bundle decisions about what to compress, when to compress, and how much to remove into fixed policies. A systematic characterization is needed to disentangle these decisions and reveal how each affects task success and execution cost. We systematically vary these decisions across three open-weight models on SWE-bench Verified and Terminal-Bench 1.0. Across nearly 35,000 agent runs, we measure task success, token use, end-to-end latency, and estimated cost. We find that fewer tokens need not mean faster or cheaper execution: on Terminal-Bench with Qwen, policies using roughly one-third as many tokens can take 20-80% longer than the uncompressed agent. Policies with similar overall success can solve different tasks, while the same policy can perform quite differently across models. Our results motivate evaluating compression by its effects on agent execution and tailoring policies to the task, model, and workload.