MechBench: Can AI Scientific Agents Discover Mechanisms Beyond Phenomenal Laws?

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation of existing scientific agent benchmarks, which evaluate only the recovery of phenomenological laws while neglecting the discovery of underlying mechanisms. We propose a novel evaluation framework that explicitly decouples phenomena from mechanisms. By employing controlled model perturbations and mechanistic probing to rule out memorization effects, combined with symbolic regression and indistinguishability filtering, this work systematically assesses agents' generalization in mechanistic reasoning. Our results reveal a pronounced gap between phenomenological and mechanistic recovery, demonstrating that current AI scientific agents suffer from severe generalization bottlenecks in deep mechanistic reasoning. This research establishes a new paradigm for future evaluation systems in scientific artificial intelligence.
πŸ“ Abstract
Scientific discovery requires not only recovering mathematical laws that describe observable behavior, but also identifying the mechanisms that generate them. Existing benchmarks for symbolic regression and scientific agents primarily evaluate phenomenal-law recovery, leaving mechanism discovery largely untested. We introduce MechBench, a benchmark that explicitly separates these two capabilities. Each task is defined by a mechanistic model, a structured set of scientifically meaningful relations whose joint consequences entail an observable phenomenal law, while agents receive only observational data and scientific context. We evaluate mechanism recovery through mechanism probes, which query internal scientific consequences that cannot be inferred from the phenomenal law alone. To reduce reliance on memorized textbook mechanisms, we construct unfamiliar variants through controlled, scientifically interpretable mutations of canonical mechanisms, and screen for mechanistic indistinguishability to exclude ambiguous instances admitting comparable competing mechanisms. Experiments across representative scientific agents reveal a substantial phenomenal--mechanism recovery gap: for Codex with GPT-5.6-sol, phenomenal-law accuracy reaches 35.00% on the Core-set while mechanism accuracy is only 13.75%, with mechanism recovery failing in 64.29% of cases where the phenomenal law is correctly recovered. The gap widens as mechanisms become increasingly mutated, and even providing the correct phenomenal law leaves mechanism recovery below 50%. These results reveal a substantial generalization gap in mechanistic reasoning and establish mechanism discovery as a distinct challenge beyond recovering observable scientific laws.
Problem

Research questions and friction points this paper is trying to address.

Mechanism Discovery
Scientific Agents
Phenomenal Laws
Symbolic Regression
Mechanistic Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mechanism Discovery
MechBench
Mechanism Probes
Scientific Agents
Symbolic Regression
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Zihan Yu
Zihan Yu
MSc Student at Imperial College London
Computer VisionMedical AI
J
Jiadong Zhang
Institute of Automation, Chinese Academy of Sciences, Beijing, China; Beijing Zhongguancun Academy, Beijing, China
J
Jialin Cheng
Xi’an Jiaotong University, Xi’an, China
Jingtao Ding
Jingtao Ding
Tsinghua University
Spatio-temporal Data MiningComplex NetworksSynthetic DataRecommender Systems
Y
Yong Li
Department of Electronic Engineering, BNRist, Tsinghua University, Beijing, China