🤖 AI Summary
Current large language models struggle to perform reliable, evidence-constrained scientific reasoning based on experimental results. To address this limitation, this work introduces SEE, the first multimodal evaluation benchmark tailored to real-world experimental science, which integrates peer-reviewed literature and associated experimental data from chemistry, biology, and materials science. The benchmark further incorporates a tool-augmented visual agent evaluation paradigm. Using an expert-curated question set, we evaluate scientific reasoning capabilities across 19 models and find that general-purpose models outperform domain-specialized ones. Although integrating external tools improves accuracy from 48.7% to 52.7%, overall performance remains limited, highlighting a critical challenge: tool use does not necessarily enhance the reliability of scientific reasoning.
📝 Abstract
Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science. Here we introduce Science Edge Evaluation (SEE), a multimodal benchmark of expert-curated questions grounded in peer-reviewed literature and experimental practice in chemistry, biology, and materials science. Evaluation of 19 multimodal large language models (MLLMs) shows that even the best-performing model reaches only 48.7% accuracy. Moreover, general-purpose models outperform science-specialized models on average. In the visual-agent evaluation, the use of tools increases the best accuracy to 52.7%. Tool use can expand the information available to models, but more information does not necessarily lead to reliable scientific reasoning. The key challenge is whether models can manage tool-derived information within the boundaries of the original experimental evidence. Together, these findings reveal that current MLLMs still cannot reliably make justified and evidence-bounded inferences from experimental results, which is an essential capability in real scientific discovery. Bridging this gap requires MLLMs to transition from explaining established scientific concepts to deriving novel and evidence-based insights from experimental data.