AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation

📅 2026-07-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study evaluates the capabilities and potential biases of large language models (LLMs) in planning and peer-reviewing research proposals in physics, astrophysics, and cosmology. For the first time, AI systems—including ChatGPT, Claude, and DeepSeek—were systematically integrated as both proposal authors and peer reviewers alongside human researchers and experts in a double-blind assessment. Human reviewers judged AI- and human-authored proposals to be of comparable quality and could identify the author type with only ~75% accuracy. In contrast, AI reviewers (Claude Opus 4.8 and ChatGPT Pro 5.5) exhibited a strong systematic preference for AI-generated content and identified its origin with 100% accuracy. These findings highlight both the innovative potential and significant risks of deploying LLMs as autonomous agents in scientific peer review.
📝 Abstract
We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were independently generated for eight expert-conceived research projects in physics, astrophysics, and cosmology by human researchers and three contemporary LLMs (ChatGPT, Claude, and DeepSeek; mid-2025 models, used with their default tool access). The resulting 32 proposals were blindly evaluated by four human reviewers and two newer frontier LLMs (Claude Opus 4.8 and ChatGPT Pro 5.5) using a four-aspect evaluation rubric. Reviewers were also asked to identify whether each proposal was written by a human or an AI. Human reviewers rated human- and AI-written proposals similarly overall, whereas both AI reviewers scored AI-written proposals about one point higher (on a five-point scale) than human-written proposals. Human reviewers correctly identified human- and AI-written proposals 72% and 79% of the time, respectively, while both AI reviewers correctly classified all 32 proposals (100%). These results suggest that current LLMs can produce project plans comparable to human-written ones in the eyes of human reviewers, but that AI reviewers show a systematic preference for AI-generated proposals. Our results suggest caution when deploying LLMs widely in proposal preparation and evaluation.
Problem

Research questions and friction points this paper is trying to address.

large language models
scientific project planning
proposal evaluation
AI-generated content
human-AI comparison
Innovation

Methods, ideas, or system contributions that make the work stand out.

large language models
scientific proposal evaluation
AI-human comparison
project planning
bias in AI review