🤖 AI Summary
This study addresses the module redundancy in agent frameworks caused by iterative patching, which increases token overhead without improving performance. To this end, this work introduces the concept of saliency from neural network pruning into agent framework optimization for the first time, proposing a saliency-based pruning strategy. Specifically, the method computes saliency scores via single-module ablation experiments and conducts quantitative evaluations that jointly consider task performance and token cost to precisely remove inefficient components. Empirical results demonstrate that existing frameworks are generally highly redundant; after substantial pruning, they can maintain comparable performance while significantly reducing computational overhead. By effectively balancing performance and efficiency, this approach offers a novel perspective for the design of agent frameworks.
📝 Abstract
Agent harnesses are systems that coordinate model calls, tool use, and task execution to help large language models complete complex tasks. To meet task requirements and address failures, these systems are often iteratively refined by amending and patching their instructions, tools, and workflows, continuously increasing harness complexity. It is therefore unclear whether some resulting harness modules are redundant, introducing substantial token overhead with little, if any, performance gain. Inspired by neural network pruning, in this paper, we study harness pruning as a means of striking a better balance between task performance and token cost. We propose SHarP (Saliency-based Harness Pruning), a simple yet effective pruning strategy based on the saliency of each harness module with respect to performance and efficiency. Specifically, we first identify tools, instructions, and supporting mechanisms as components that can be individually disabled. We then estimate the saliency of each module by ablating it and assessing its task performance and token cost relative to the full set of single-module ablations. Modules with the smallest contribution to performance or largest computational overhead are subsequently pruned. Our evaluation across various harnesses on held-out validation sets reveals a surprising finding: most harnesses that we studied are highly redundant and can maintain comparable performance and efficiency even after a substantial portion of their modules are pruned. Our pruning approach and empirical findings provide new perspectives on agent harness design and optimization.