Unable to forget: Proactive lnterference Reveals Working Memory Limits in LLMs Beyond Context Length

📅 2025-06-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study identifies a working-memory-like bottleneck in large language models (LLMs): when updating semantically related key-value sequences, prior information induces proactive interference, causing retrieval accuracy for the most recent item to decay logarithmically-linearly to zero with increasing interference count—beyond standard context-length limitations. Method: Drawing on cognitive science, we introduce the proactive interference (PI) paradigm to LLM evaluation, proposing the PI-LLM benchmark. It features streaming key-value injection, controllable interference intensity, and multi-condition prompt ablations. Contribution/Results: We demonstrate that LLMs lack dynamic content suppression capability; instruction-based prompting yields only marginal mitigation. These findings directly challenge the prevailing “longer context implies stronger retrieval” assumption. Crucially, we establish intrinsic suppression deficits as a fundamental limitation of current LLMs and provide a reproducible, cognitively inspired evaluation framework for diagnosing such deficits.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsCognitive Modeling & Cognitive Systems: Conceptual Inference and Reasoning

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
📝 Abstract
Information retrieval in Large Language Models (LLMs) is increasingly recognized as intertwined with generation capabilities rather than mere lookup. While longer contexts are often assumed to improve retrieval, the effects of intra-context interference remain understudied. To address this, we adapt the proactive interference (PI) paradigm from cognitive science, where earlier information disrupts recall of newer updates. In humans, susceptibility to such interference is inversely linked to working memory capacity. We introduce PI-LLM, an evaluation that sequentially streams semantically related key-value updates and queries only the final values. Although these final values are clearly positioned just before the query, LLM retrieval accuracy declines log-linearly toward zero as interference accumulates; errors arise from retrieving previously overwritten values. Attempts to mitigate interference via prompt engineering (e.g., instructing models to ignore earlier input) yield limited success. These findings reveal a fundamental constraint on LLMs' ability to disentangle interference and flexibly manipulate information, suggesting a working memory bottleneck beyond mere context access. This calls for approaches that strengthen models' ability to suppress irrelevant content during retrieval.
Problem

Research questions and friction points this paper is trying to address.

Investigates interference effects on LLM memory retrieval
Evaluates LLMs' working memory limits beyond context length
Explores mitigation of retrieval errors from overwritten data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adapts proactive interference paradigm from cognitive science
Introduces PI-LLM for evaluating retrieval accuracy
Highlights need to suppress irrelevant content
🔎 Similar Papers
No similar papers found.