Does an Agent's History Tell You When Compaction Will Hurt? A Modest, Bounded Effect on the TRACE Paired-Replay Corpus

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether the historical behaviors of long-horizon agents can predict performance degradation induced by context compression. Leveraging the TRACE corpus, the authors replay compression boundaries to analyze how prefix states influence subsequent actions. Through a paired replay mechanism, AUROC evaluation, and internal comparative experiments, the work reveals the predictive limitations of historical signals and proposes interpretable triggers to balance information retention with damage mitigation. The results demonstrate that an optimal trigger successfully avoids 21% of harmful compression boundaries while preserving 84% of beneficial opportunities. These findings provide empirical evidence and novel insights for understanding risk prediction in agent context compression.
📝 Abstract
Many long-horizon agents compact their context on a global rule, usually a token budget, blind to what the agent was doing. We ask whether the agent's recent behaviour predicts when a compaction will hurt. TRACE's public corpus of 590 harness-triggered AppWorld compaction boundaries replays each boundary from a re-executed prefix state under the pre-compaction context and under the summary, and records the burden of the next actions: calls that error or repeat a call already made. We find that pre-boundary history predicts post-compaction harm only weakly. An internally prespecified contrast by prefix placement is a wide null, and the naive "has-written" label behind it turns out to measure trajectory phase. The best extension-protocol trigger reaches held-out AUROC 0.66 (0.64 on the replicate's own label) against a same-boundary replicate of 0.72; the best frozen, interpretable trigger avoids 21% of harmful (positive-burden) boundaries while keeping 84% of compaction opportunities, and exceeds the random-rule expectation on count but not on burden mass (a post hoc comparison). Whether the best trigger beats a token-budget rule at matched retention cannot be evaluated on the release. We state what corpora should ship to answer it.
Problem

Research questions and friction points this paper is trying to address.

context compaction
long-horizon agents
compaction harm prediction
agent history
token budget
Innovation

Methods, ideas, or system contributions that make the work stand out.

context compaction
long-horizon agents
trajectory replay
predictive triggering
interpretable rules