Continuous Context Management

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the context redundancy and prohibitive costs incurred when large language model agents retain complete interaction histories. To mitigate this, we propose a continuous context management paradigm that introduces a novel per-turn immediate compression mechanism, replacing full transcription with dynamic memory updates. Furthermore, we present a privileged full-history distillation technique that enables dense supervision without requiring an independent teacher model, jointly optimizing the compression strategy via GRPO-based reinforcement learning. Experimental results demonstrate that our approach significantly reduces input consumption and prompt volume while outperforming baselines on benchmarks such as WebShop, thereby validating the feasibility of low-context reasoning for LLM agents.
📝 Abstract
Long-horizon large language model (LLM) agents commonly retain their complete interaction history until compaction is triggered at a predefined threshold. We study Continuous Context Management (CCM), which performs compaction at every turn to prevent interaction history from accumulating in the active prompt. At each turn, a CCM agent emits an updated memory together with an environment action; its next prompt contains the original task, retained memory, and newest observation rather than the complete transcript. We first evaluate CCM without fine-tuning on TerminalBench-2 using Claude Sonnet 4.6, Claude Opus 4.6, GLM-5, and Kimi K3. CCM substantially reduces cumulative input usage and active-prompt size, although it lowers task success for most models while preserving performance for Kimi K3. We use GRPO with privileged full-history distillation to improve CCM in open-weight models. A frozen copy of the student's initial model scores each sampled student action under the complete history reconstructed from that student's rollout, providing dense action-token supervision without a separate teacher rollout or reference solution. On WebShop, this objective substantially improves CCM over GRPO at both evaluated model scales and surpasses full-history GRPO for Qwen3-4B-Instruct, though not for Qwen3-8B. On Endless Terminals, the augmented method provides a modest improvement over GRPO, with both CCM policies outperforming the untrained full-history baseline. These results demonstrate that CCM is a viable inference paradigm for agents operating with substantially reduced retained context and that its performance can be improved through reinforcement learning with privileged full-history distillation.
Problem

Research questions and friction points this paper is trying to address.

Continuous Context Management
Long-horizon LLM agents
Context compaction
Prompt size reduction
Reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Continuous Context Management
Privileged Full-History Distillation
GRPO
Long-horizon LLM Agents
Context Compaction