🤖 AI Summary
This study addresses the vulnerability of large language models to "stale bindings" during in-context updates, wherein outdated values persist in generated outputs. We reveal that this failure stems from attention drift: although probes can recover representations of updated values, the generation process remains dominated by stale information due to ineffective attentional competition. To investigate this phenomenon, we introduce the CICM benchmark alongside a theoretical analysis framework based on single-layer Transformers, which precisely identifies the bottleneck in attentional competition. Building on these insights, we propose a training-free inference-time intervention technique that redirects attention to mitigate this failure. Our approach significantly reduces context-update errors across diverse model architectures while preserving original task accuracy without degradation. This work offers an efficient and interpretable solution for dynamic information updating in large language models.
📝 Abstract
As preferences, goals, and facts change, LLM agents must use the current state while earlier versions remain in context. Yet they can answer with an old value of the same variable, a failure that we call stale binding. To study when models use outdated information and why, we introduce Controlled In-Context Memory (CICM), a benchmark for tracking and using updated information in conversations and agent logs. We observe that even frontier reasoning models can fail to recover the current state. We find that in open-source models probes can still recover the updated value when the model answers with an old one, pointing to a failure to select information that remains available. Component tests in Qwen and Pythia identify a mechanism for this selection failure: attention drift, where attention favors old values over the current one when producing an answer. We study a one-layer transformer to mathematically understand how this phenomenon happens: when attention scores are similar, several old values can together receive more attention than the current value. Guided by this explanation, we redirect attention toward the current value without further training. When the current value is requested directly, adjusting this intervention for each input corrects most old-value errors across various model families while preserving nearly all initially correct answers. Reliable context management therefore requires more than remembering updated information: models must use it to guide their answers.