🤖 AI Summary
This study investigates whether AI agents can autonomously discover and interpret novel scientific mechanisms from observational data in long-horizon experiments. To this end, it proposes a three-axis evaluation framework centered on deriving scientific insights from underlying mechanisms and constructs an interdisciplinary benchmark spanning six domains, including neuroscience, comprising 306 expert-validated tasks for expected scientific insights. The findings reveal that while current agents surpass human baselines in predictive accuracy, they exhibit substantial deficiencies in inferring deeper scientific insights. These results delineate critical directions for advancing the capabilities of future scientific research agents.
📝 Abstract
When Isaac Newton discovered the law of gravitation, he did so through an iterative process of analyzing observed data such as planetary patterns, finding the underlying mechanisms by describing patterns in mathematical equations, and refining his theory against the Moon's orbit, revealing the startling insight that the same force governs both falling apples and orbiting planets. Would it be possible for AI agents to make similar discoveries? To measure this ability, we introduce EurekaBench, a cross-domain benchmark that tests AI agents' ability to conduct long-horizon experiments and discover mechanisms that explain observations. We evaluate these mechanisms by the scientific insights that can be derived from them. EurekaBench contains an expert-verified set of 26 long-horizon tasks across neuroscience, computer science, chemistry, astrophysics, geophysics, and plasma physics, with a total of 306 scientific insights that the discovered mechanisms are expected to support. Our evaluation framework tests three axes of scientific discovery: agents' ability to follow known scientific constraints, the predictive accuracy of the discovered mechanisms, and whether these mechanisms yield scientific insights or inform future research. Our results show that current AI agents often overly fixate on predictive accuracy optimization, surpassing human scientists, while falling substantially short in deriving scientific insights.