Do Large Language Models know Which Published Articles have been Retracted?

📅 2026-04-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the risk that large language models (LLMs), when operating without online retrieval support, may propagate misinformation by failing to recognize retracted scientific literature. It presents the first systematic evaluation of open-source LLMs—including GPT OSS 120B, Gemma 3 27B, and DeepSeek R1 72B—in their ability to identify high-impact retracted articles based solely on titles and abstracts, benchmarked against a large-scale dataset of non-retracted papers. Results reveal that these models fail to detect retracted articles more than 80% of the time, while maintaining an extremely low false-positive rate (<0.2%) for valid publications. These findings highlight a critical limitation in current LLMs’ capacity to assess scholarly credibility and underscore the necessity of incorporating retraction status either into model training data or via retrieval-augmented mechanisms to mitigate the dissemination of discredited research.

Technology Category

Natural Language Processing: Fact-Checking / Misinformation Detection (NLP Focus)Machine Learning: Large Multimodal Models (LMMs)Data Mining & Knowledge Management: Conversational Systems for Recommendation & Retrieval

Application Category

Search and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Large Language Models (LLMs) can be helpful for literature search and summarisation, but retracted articles can confuse them. This article asks three open weights (offline) LLMs whether 161 high profile retracted articles had been retracted, performing a similar check for a benchmark multidisciplinary set of 34,070 non-retracted articles. Based on titles and abstracts, in over 80% of cases the LLMs claimed that a retracted article had not been retracted (GPT OSS 120B: 82%; Gemma 3 27B: 84%; DeepSeek R1 72B: 88%). The reasons given for a correct retraction declaration were often wrong, even if detailed. This confirms that LLMs have little ability to distinguish between valid and retracted studies, unless they are allowed to, and do, check online. For the benchmark test, there were only 55 false retraction claims from 34,070 non-retracted full text articles, and 28 false claims when only the title and abstract were entered, suggesting that there is only a small chance that LLMs discount valid studies. When retractions are erroneously claimed, this does not seem to be due to mistakes in the article. Overall, the results give new reasons to be cautious about LLM claims about academic findings.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Retracted Articles
Literature Search
Academic Integrity
Scientific Misinformation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
retracted articles
scientific literature evaluation
offline LLM limitations
academic integrity
🔎 Similar Papers
No similar papers found.