Negation-Induced Forgetting in LLMs

📅 2025-02-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study empirically investigates whether large language models (LLMs) exhibit “negation-induced forgetting” (NIF)—a cognitive phenomenon wherein negating incorrect attributes impairs subsequent recall of the target object—previously documented in human memory research. Method: Adapting Zang et al.’s behavioral paradigm, we conducted controlled prompting and recall experiments on ChatGPT-3.5, GPT-4o-mini, and Llama3-70b-instruct, systematically varying negation cues and measuring recall accuracy. Contribution/Results: We report the first empirical evidence of NIF in LLMs: ChatGPT-3.5 shows a statistically significant NIF effect; GPT-4o-mini exhibits marginally significant NIF; and Llama3-70b-instruct shows no reliable NIF. These findings demonstrate that certain LLMs instantiate human-like memory biases, offering novel empirical grounding for understanding their internal memory dynamics. Moreover, this work advances human-inspired cognitive modeling of LLMs by bridging computational linguistics with cognitive psychology, suggesting that memory phenomena observed in humans may partially generalize to artificial language systems under specific architectural and training conditions.

Technology Category

Natural Language Processing: (Large) Language ModelsMachine Learning: Large Multimodal Models (LMMs)Cognitive Modeling & Cognitive Systems: Conceptual Inference and Reasoning

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
The study explores whether Large Language Models (LLMs) exhibit negation-induced forgetting (NIF), a cognitive phenomenon observed in humans where negating incorrect attributes of an object or event leads to diminished recall of this object or event compared to affirming correct attributes (Mayo et al., 2014; Zang et al., 2023). We adapted Zang et al. (2023) experimental framework to test this effect in ChatGPT-3.5, GPT-4o mini and Llama3-70b-instruct. Our results show that ChatGPT-3.5 exhibits NIF, with negated information being less likely to be recalled than affirmed information. GPT-4o-mini showed a marginally significant NIF effect, while LLaMA-3-70B did not exhibit NIF. The findings provide initial evidence of negation-induced forgetting in some LLMs, suggesting that similar cognitive biases may emerge in these models. This work is a preliminary step in understanding how memory-related phenomena manifest in LLMs.
Problem

Research questions and friction points this paper is trying to address.

Explores negation-induced forgetting in LLMs
Tests NIF effect in ChatGPT-3.5, GPT-4o mini, Llama3-70b
Provides evidence of cognitive biases in some LLMs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tested negation-induced forgetting in LLMs
Adapted experimental framework for LLMs
Identified NIF in ChatGPT-3.5
🔎 Similar Papers
No similar papers found.
F
Francesca Capuano
Department of Psychology, University of Tübingen
E
Ellen Boschert
Department of Psychology, University of Tübingen
Barbara Kaup
Barbara Kaup
Universität Tübingen FB Psychologie