Are We Measuring Scientific Intelligence? Rethinking the Evaluation of AI Scientists

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a critical limitation in existing AI-for-science evaluation benchmarks, which assume correct answers derive from provided data while overlooking distortions caused by models relying on prior memorization or elimination strategies. We propose an "evidence-anchored accuracy" metric that applies retention, withdrawal, and reversal perturbations to data evidence, combined with preregistered statistical validation and multi-model agent comparative experiments, to rigorously determine whether models genuinely reason from data. Our findings reveal that Claude achieves 83% accuracy on single-cell tasks without data access, yet attains only 41% evidence-anchored accuracy, demonstrating that high scores reflect memorization rather than reasoning. This work critically re-examines scientific intelligence evaluation frameworks, proving that conventional scoring metrics fail to distinguish between evidence-driven reasoning and reliance on prior knowledge.
📝 Abstract
AI agents can now carry out data-driven scientific analyses end to end, and benchmarks assess them by giving an agent a question and a dataset and scoring its final answer against a fixed key. These benchmarks assume that a correct answer was derived from the supplied data, a property we call evidence grounding. However, an agent can also reach the key from prior knowledge or by ruling out the other options, and a score based on a single run cannot tell these cases apart. We show how to test this assumption and find that it often fails. For each question, we build versions of its data files in which the evidence for the answer is left intact, withdrawn or reversed, check each edit with a pre-registered reference statistic, and run the same agent on every version. We then measure evidence-grounded accuracy, which credits a correct answer only if the agent also responds when the evidence is withdrawn and follows it when it is reversed. We evaluate three agent scaffolds and five models on 18 single-cell questions from BAISBench and four synthetic problems from GeneBench-Pro. On the single-cell questions, the Claude agents are 95% accurate and answer 83% correctly without any data, but their evidence-grounded accuracy is only 41%. Hiding gene names raises the share of runs that follow reversed evidence from 58% to 93% on ten gene tasks, suggesting that prior knowledge competes with the supplied data. The benchmark score and LLM judges can also reward answers that ignore the changed evidence. Measuring scientific intelligence rather than recall therefore requires checking whether answers follow the evidence and whether scores reward them for it.
Problem

Research questions and friction points this paper is trying to address.

AI scientist evaluation
evidence grounding
scientific intelligence
benchmark reliability
prior knowledge leakage
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evidence Grounding
AI Scientist Evaluation
Counterfactual Data Perturbation
Scientific Intelligence
Benchmark Reliability
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kate Zhang
School of Computer Science, Carnegie Mellon University, Pittsburgh, PA 15213
Yuante Li
Yuante Li
Carnegie Mellon University
AI ScientistMulti-Agent SystemLarge Language ModelsData MiningAI For Finance