Downstream-Aware Context Selection for Online In-Context Reinforcement Learning

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive token costs arising from the accumulation of historical interactions in online in-context reinforcement learning. To mitigate this, it proposes a bounded history management framework that dynamically guides history selection and budget allocation by predicting the downstream effects of removing specific interactions. Using the full rolling history as a reference, the framework employs task-dependent predictors to evaluate deletion impacts and establish prioritization rankings, subsequently applying a shared criterion to achieve efficient history compression. Experimental evaluations in driving and ScienceWorld environments demonstrate that this approach reduces token consumption by 25% and 52%, respectively, while maintaining task performance comparable to the original system.
📝 Abstract
In-context reinforcement learning (ICRL) enables large language model agents to adapt to new environments using their interaction history without updating model parameters. However, repeatedly conditioning on growing histories can lead to substantial token cost. We propose a bounded-history context-management framework that predicts the task-dependent downstream effect of removing historical interactions to guide history selection and determine a decision-dependent context budget. Formally, our framework uses the full rolling history as a reference. The predictor evaluates removal effects, defines a deletion ordering, and applies a shared selection criterion to determine how much history to retain at each decision. We evaluate the method in closed-loop SUMO driving under held-out in-distribution, unseen-domain, and unseen-route settings, and in ScienceWorld under a continual ICRL protocol. Relative to a baseline using the full context, our method reduces total token usage by 25.7%, 25.8%, and 23.2% across the three driving settings while maintaining comparable closed-loop driving performance. In ScienceWorld, it reduces total token usage by 52.1% compared to full context and uses 30.2% and 37.8% fewer tokens than the Recent and Similarity baselines, respectively, while maintaining performance.
Problem

Research questions and friction points this paper is trying to address.

In-context reinforcement learning
token cost
context management
history selection
large language model agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-Context Reinforcement Learning
Context Selection
Bounded-History Framework
Downstream-Aware Prediction
Token Efficiency