TAP: Efficient Long-Horizon Agent Pruning via Trajectory-Anchored Recovery

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive inference costs of long-horizon agents and the severe performance degradation caused by existing pruning methods. We propose the first structured pruning framework tailored for reinforcement learning-trained agents, which innovatively couples structural pruning with online policy recovery. Methodologically, compression is achieved via gradient-based iterative channel re-scoring, while distribution shift and error accumulation are effectively mitigated through teacher-trajectory-anchored interaction, frozen dense teacher-supervised prefix generation, and online distillation. Experimental results demonstrate that after removing 60% of FFN channels, the proposed method preserves success rates of 99.2% on ALFWorld and 88.0% on WebShop while reducing GPU inference time by approximately 20%, thereby achieving highly efficient model compression without compromising task performance.
📝 Abstract
Emerging long-horizon agentic tasks require repeated model calls, worsening the inference cost of already-costly language models. While narrow agentic tasks suggest potential for aggressive model pruning without performance drop, empirical results show existing methods proposed for question answering tasks severely degrade task performance when applied to agentic models. We trace this failure to two decisions: what to prune and how to recover. For pruning, one-shot importance estimates fail to track how the pruned model adapts. For recovery, offline distillation covers only teacher prefixes, while full-trajectory on-policy distillation causes student errors to compound across turns. In this work, we propose Trajectory-Anchored Pruning (TAP), the first structural pruning framework for reinforcement learning (RL)-trained agents. TAP couples structural pruning with efficient on-policy recovery, anchoring interactions to teacher trajectories while allowing the student to generate each reasoning-action response. A frozen dense teacher supervises the student's response prefixes, addressing within-response training-inference mismatch while preventing student-induced deviations from propagating across training turns. Instead of one-shot pruning, TAP re-scores channels using gradients of the recovery objective on the recovered student, connecting iterative channel selection to the evolving policy. With 60% of FFN channels removed, TAP retains 99.2% and 88.0% of the dense 7B agents' task success rates on ALFWorld and WebShop, respectively, while reducing GPU time per successful task by approximately 22% and 17%. These results demonstrate effective structural compression of long-horizon agents under a limited recovery budget.
Problem

Research questions and friction points this paper is trying to address.

Long-horizon agents
Structural pruning
Inference cost
Agent compression
Recovery distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Trajectory-Anchored Pruning
Structural Pruning
Reinforcement Learning Agents
On-policy Recovery
Long-Horizon Tasks
🔎 Similar Papers
No similar papers found.