AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review

📅 2026-07-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study systematically evaluates the reliability of large language models in assisting with literature reviews in physics, astrophysics, and cosmology. Through eight expert-designed controlled experiments, it comparatively analyzes retrieval outcomes from human experts and leading models—including ChatGPT-4o, Gemini, and the 2026 ChatGPT Pro 5.5—revealing a human–model citation overlap of less than 6% in these highly specialized domains. The work introduces a novel taxonomy of hallucinations, distinguishing between “fabricated references” and “metadata mismatches.” Results indicate that models released in 2025 generated, on average, 3% fabricated references and 64% metadata errors, whereas ChatGPT Pro 5.5 achieved zero errors in a single-task evaluation, demonstrating substantially improved accuracy and trustworthiness.
📝 Abstract
We investigate how well large language models (LLMs) can assist with literature reviews for scientific research. We perform a controlled study of eight expert-conceived research projects across the areas of physics, astrophysics, and cosmology. Each project has a defined background and goal, and human experts and AI prompters are asked to perform identical literature review tasks in parallel. We compare the relevant literature selected by humans with that selected by mid-2025 LLMs (ChatGPT-4o, ChatGPT Deep Research, and Gemini). We find the overlap between human- and AI-selected references to be small ($<$6\%), indicating that AI models do not yet reproduce a competent expert search on their own, though they have the potential to complement literature searches by humans. We then assess the reliability and completeness of AI-generated candidate references, distinguishing two types of hallucination: fabrications (references to nonexistent papers) and metadata mismatches (real papers with one or more incorrect fields). We find that while fabricated references make up 3\% of the AI-generated references, 64\% are real papers with at least one incorrect field (title, author, year, journal, DOI, or link), indicating that the mid-2025 models require systematic verification. However, the performance is significantly improved for the 2026 model ChatGPT Pro 5.5, with a single-project test showing zero fabrication or metadata mismatches.
Problem

Research questions and friction points this paper is trying to address.

literature review
large language models
scientific research
hallucination
reference accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

large language models
literature review
hallucination
metadata accuracy
scientific research assistance
🔎 Similar Papers
No similar papers found.
A
Anamaria Hell
K
Kateryna Vovk
V
Veena Krishnaraj
J
Jia Liu
K
Kosuke Aizawa
A
Adrian E. Bayer
L
Linda Blot
J
Jessica Cowell
S
Suyog Garg
J
Jonathan Grée
B
Ben Horowitz
M
Masaya Ichikawa
K
Kanyuni Iemoto
K
Keigo Kondo
Z
Zacharie Lorsin
Kevin McCarthy
Kevin McCarthy
Former Research Fellow, INSIGHT, University College Dublin, Eire
PersonalizationRecommender SystemsInformation RetrievalMachine Learning
J
Jamie Robinson
M
Miguel Ruiz-Granda
L
Leander Thiele
I
Ievgen Vovk
M
Mingshen Zhou