From Scientific Observations to Mechanisms: Benchmarking Hypothesis Generation by AI Scientists

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of effective evaluation frameworks for assessing the ability of AI scientists to translate empirical data into mechanistic hypotheses. To this end, we introduce MechHypoBench, the first benchmark for mechanistic hypothesis generation, which integrates multidisciplinary literature with large-scale real-world data. Furthermore, we propose an open-ended hypothesis evaluation framework grounded in consequences under hidden conditions. By leveraging cross-disciplinary knowledge integration, open hypothesis generation, and causal inference evaluation techniques, this work systematically examines the capacity of AI systems to distill underlying mechanisms from observations. Experimental results demonstrate that a significant gap persists between the hypotheses generated by current general-purpose agents or AI scientists and the ground-truth mechanisms, thereby highlighting both the challenges and the potential of this research direction.
📝 Abstract
Data-driven mechanistic hypotheses are essential to scientific discovery because they explain how underlying processes produce observed phenomena. AI agents and AI scientists increasingly support scientific data analysis. However, their ability to turn empirical findings into mechanistic hypotheses remains insufficiently examined. To address this gap, we introduce MechHypoBench, the first benchmark for evaluating whether AI agents and AI scientists can generate such hypotheses from empirical data. It combines paper-derived mechanisms from 14 scientific fields with real-world datasets containing 17.98 million records. The construction retains the observational complexity of empirical data while providing a specified underlying mechanism. Agents analyze the observations and propose open-form hypotheses. We develop an evaluation framework that assesses open-form mechanistic hypotheses through their consequences under withheld conditions. Experiments with general agents and AI scientists reveal a substantial gap between generated hypotheses and the underlying mechanisms.
Problem

Research questions and friction points this paper is trying to address.

Mechanistic Hypothesis Generation
AI Scientists
Benchmarking
Scientific Discovery
Empirical Data
Innovation

Methods, ideas, or system contributions that make the work stand out.

MechHypoBench
Mechanistic Hypotheses
AI Scientists
Evaluation Framework
Benchmarking
🔎 Similar Papers
X
Xiaxun Xie
Computer Network Information Center, Chinese Academy of Sciences
Q
Qingqing Long
Computer Network Information Center, Chinese Academy of Sciences
M
Meng Xiao
National University of Singapore
W
Wei Ju
Sichuan University
Yuanchun Zhou
Yuanchun Zhou
Computer Network Information Center,CAS
Data MiningBig Data Analysis
Xuezhi Wang
Xuezhi Wang
Research Scientist, Google DeepMind
Machine LearningNatural Language Processing
H
Hengshu Zhu
Computer Network Information Center, Chinese Academy of Sciences