ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of context window overflow, high computational costs, and irreversible information loss in long-horizon LLM agents caused by cumulative contexts. To this end, we propose ReFold, a training-free reversible context folding framework. By leveraging chunk-based rendering, inter-turn redundancy detection, and a compression algorithm utilizing stubs and single-line notes, ReFold reversibly folds redundant interaction histories into a compact representation, enabling plug-and-play integration with existing agent frameworks. Experimental results demonstrate that ReFold reduces token consumption by 2.5×, halves the KV cache size, accelerates inference by 1.7×, and lowers overall costs by 3.4×, all while preserving task success rates without degradation.
📝 Abstract
Long-horizon LLM agents act on an append-only interaction history that is re-sent to the model at every step, so the context and its cost grow with steps until the sessions exceed the context window. Existing methods manage the context through context requirement prediction, relying on additional model calls, heuristic rules, or trained policies. However, these predictive approaches introduce runtime overhead, invalidate prefix caches, and permanently discard content with no guarantee of recovery. To overcome these limitations, we introduce ReFold: a training-free rendering layer that preserves the underlying interaction history while compressing only the model's rendered context. It removes two kinds of inter-turn redundancy without an auxiliary predictor: content an earlier turn already displayed, replaced by a stub, and turns the agent itself reports finished, folded into a one-line note. Both operators use chunked rendering, rewriting the cached prefix once every few steps rather than at every step. Every removal is strictly reversible, a wrong removal costs one restore from the history rather than permanent content loss. Because it operates at the rendering layer, ReFold is plug-and-play across standard ReAct-style harnesses. Evaluations across five long-horizon benchmarks and two frontier LLMs demonstrate that ReFold reduces token consumption by up to 2.5x and halves the KV-cache memory per session without degrading task success rates. Under capped context budgets, it avoids up to 92% of forced compactions. Under concurrent serving workloads, it reduces request queuing delays by up to 100%, accelerating inference by up to 1.7x, while cutting inference costs by up to 3.4x.
Problem

Research questions and friction points this paper is trying to address.

long-horizon agents
context window limit
token consumption
prefix cache invalidation
irreversible content loss
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-Free
Reversible Context Folding
Rendering Layer
Chunked Rendering
Long-Horizon Agents