🤖 AI Summary
This study addresses a critical limitation in existing AI-for-science evaluation benchmarks, which assume correct answers derive from provided data while overlooking distortions caused by models relying on prior memorization or elimination strategies. We propose an "evidence-anchored accuracy" metric that applies retention, withdrawal, and reversal perturbations to data evidence, combined with preregistered statistical validation and multi-model agent comparative experiments, to rigorously determine whether models genuinely reason from data. Our findings reveal that Claude achieves 83% accuracy on single-cell tasks without data access, yet attains only 41% evidence-anchored accuracy, demonstrating that high scores reflect memorization rather than reasoning. This work critically re-examines scientific intelligence evaluation frameworks, proving that conventional scoring metrics fail to distinguish between evidence-driven reasoning and reliance on prior knowledge.
📝 Abstract
AI agents can now carry out data-driven scientific analyses end to end, and benchmarks assess them by giving an agent a question and a dataset and scoring its final answer against a fixed key. These benchmarks assume that a correct answer was derived from the supplied data, a property we call evidence grounding. However, an agent can also reach the key from prior knowledge or by ruling out the other options, and a score based on a single run cannot tell these cases apart. We show how to test this assumption and find that it often fails. For each question, we build versions of its data files in which the evidence for the answer is left intact, withdrawn or reversed, check each edit with a pre-registered reference statistic, and run the same agent on every version. We then measure evidence-grounded accuracy, which credits a correct answer only if the agent also responds when the evidence is withdrawn and follows it when it is reversed. We evaluate three agent scaffolds and five models on 18 single-cell questions from BAISBench and four synthetic problems from GeneBench-Pro. On the single-cell questions, the Claude agents are 95% accurate and answer 83% correctly without any data, but their evidence-grounded accuracy is only 41%. Hiding gene names raises the share of runs that follow reversed evidence from 58% to 93% on ten gene tasks, suggesting that prior knowledge competes with the supplied data. The benchmark score and LLM judges can also reward answers that ignore the changed evidence. Measuring scientific intelligence rather than recall therefore requires checking whether answers follow the evidence and whether scores reward them for it.