Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current large language models struggle to perform reliable, evidence-constrained scientific reasoning based on experimental results. To address this limitation, this work introduces SEE, the first multimodal evaluation benchmark tailored to real-world experimental science, which integrates peer-reviewed literature and associated experimental data from chemistry, biology, and materials science. The benchmark further incorporates a tool-augmented visual agent evaluation paradigm. Using an expert-curated question set, we evaluate scientific reasoning capabilities across 19 models and find that general-purpose models outperform domain-specialized ones. Although integrating external tools improves accuracy from 48.7% to 52.7%, overall performance remains limited, highlighting a critical challenge: tool use does not necessarily enhance the reliability of scientific reasoning.
📝 Abstract
Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science. Here we introduce Science Edge Evaluation (SEE), a multimodal benchmark of expert-curated questions grounded in peer-reviewed literature and experimental practice in chemistry, biology, and materials science. Evaluation of 19 multimodal large language models (MLLMs) shows that even the best-performing model reaches only 48.7% accuracy. Moreover, general-purpose models outperform science-specialized models on average. In the visual-agent evaluation, the use of tools increases the best accuracy to 52.7%. Tool use can expand the information available to models, but more information does not necessarily lead to reliable scientific reasoning. The key challenge is whether models can manage tool-derived information within the boundaries of the original experimental evidence. Together, these findings reveal that current MLLMs still cannot reliably make justified and evidence-bounded inferences from experimental results, which is an essential capability in real scientific discovery. Bridging this gap requires MLLMs to transition from explaining established scientific concepts to deriving novel and evidence-based insights from experimental data.
Problem

Research questions and friction points this paper is trying to address.

scientific discovery
multimodal large language models
evidence-bounded reasoning
experimental data
scientific inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Science Edge Evaluation
multimodal large language models
evidence-bounded reasoning
tool-augmented agents
scientific discovery benchmark
🔎 Similar Papers
No similar papers found.
T
Taolin Han
Alibaba Group; University of Chinese Academy of Sciences
Yuchen Zhang
Yuchen Zhang
Tsinghua University
Scientific Machine Learning
J
Jinghang Wang
Alibaba Group
Y
Yun Wu
Alibaba Group
W
Wai Yuet Chiu
Alibaba Group
Z
Zhaohai Li
Qwen Team, Alibaba Group
Y
Yifei Zhang
Alibaba Group; University of Alberta
Jinxin Wang
Jinxin Wang
University of Chicago Booth School of Business
optimizationoptimal transportgenerative AI
Y
Yuhao Zhou
Alibaba Group; Zhejiang University
C
Chen Zhao
Alibaba Group; Tsinghua University
J
Jiajia Li
Alibaba Group
J
Jiaxin Li
Alibaba Group
Q
Qile Jin
Alibaba Group
K
Kewei Sun
Alibaba Group
S
Shuang Wu
Alibaba Group
W
Weiqi Zhai
Alibaba Group
R
Renquan Lv
Alibaba Group; Zhejiang University
Junchao Li
Junchao Li
Ph.D., Mechanical Engineering, University of Iowa
Machine LearningReinforcement learningFormal methodsPath PlanningRobotics
R
Ruodan Chen
Alibaba Group
Q
Qingteng Chen
Alibaba Group
Zhibo Yang
Zhibo Yang
Alibaba Group; Tsinghua University
OCRMLLMs
H
Hu Wei
Alibaba Group
L
Lin Qu
Alibaba Group
Shuai Bai
Shuai Bai
Qwen Team, Alibaba Group
Multi-Modal LearningVisual Generation
Bing Zhao
Bing Zhao
SRI International
Natural Language ProcessingMachine LearningOptimizations