Stress-testing large language model agents in a robotic chemistry laboratory

📅 2026-07-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study evaluates the reliability and adaptability of large language models in executing scientific tasks within real-world physical environments, with a focus on their ability to generate executable experimental protocols and iteratively refine them based on empirical evidence. Leveraging a robotic chemistry laboratory comprising 45 modular workstations and conducting 4,608 trials, this work extends scientific agent evaluation beyond pure reasoning to encompass physical executability and evidence-driven closed-loop adaptation, introducing a quantifiable framework for assessing deployment readiness. Results reveal that only 3.3% of generated protocols were deemed executable by expert reviewers, with the best-performing system achieving a success rate of 28.1%. Most generated workflows contained no more than 30 steps and generally lacked capabilities for workflow-level replanning or methodological reconfiguration in response to experimental outcomes.
📝 Abstract
AI is evaluated through knowledge, reasoning and plan generation, yet scientific agency requires reliable physical action and adaptation to evidence. Here, we use a robotic chemistry laboratory as a physical-world testbed to make scientific agency measurable. Its 45 modular workstations exposed as machine-readable skills enabled 4,608 trials. Only 3.3% of trials produced expert-assessed executable workflows under laboratory constraints; even the best system achieved 28.1%. Long-horizon planning remained a challenge: only three executable workflows exceeded 30 operations, although the longest contained 44. Across five rounds, experimental feedback prompted local adjustments but no workflow-level replanning or analytical-method redesign. By making physical executability and evidence-driven replanning measurable, our study provides an evidence-based assessment of deployment readiness and a diagnostic framework to guide closed-loop improvements towards physically grounded autonomous research.
Problem

Research questions and friction points this paper is trying to address.

scientific agency
physical executability
evidence-driven replanning
long-horizon planning
robotic chemistry laboratory
Innovation

Methods, ideas, or system contributions that make the work stand out.

robotic chemistry laboratory
large language model agents
physical executability
evidence-driven replanning
scientific agency
Lulu Guo
Lulu Guo
Tongji University
automotive advanced controlvehicle cybersecurityenergy management
Y
Yingkai Sun
State Key Laboratory of Precision and Intelligent Chemistry, Hefei National Research Center for Physical Sciences at the Microscale, School of Chemistry and Materials Science, University of Science and Technology of China, Hefei, China
X
Xiaobo Li
State Key Laboratory of Precision and Intelligent Chemistry, Hefei National Research Center for Physical Sciences at the Microscale, School of Chemistry and Materials Science, University of Science and Technology of China, Hefei, China; Center for Scientific Intelligence Innovation, Hefei, China
L
Luyao Ge
State Key Laboratory of Precision and Intelligent Chemistry, Hefei National Research Center for Physical Sciences at the Microscale, School of Chemistry and Materials Science, University of Science and Technology of China, Hefei, China
Ziming Wang
Ziming Wang
University of Hong Kong; University of Science and Technology of China
Embodied AIRoboticsSLAMHuman-Robot InteractionEdge Computing
Haitao Zheng
Haitao Zheng
Neubauer Professor of Computer Science, University of Chicago
Mobile ComputingSecurity and Privacy
Jingyu Li
Jingyu Li
University of Science and Technology of China
Deep LearningComputer VisionNatural Language Processing
H
Huijuan Zhang
State Key Laboratory of Precision and Intelligent Chemistry, Hefei National Research Center for Physical Sciences at the Microscale, School of Chemistry and Materials Science, University of Science and Technology of China, Hefei, China
B
Bingxu Chen
State Key Laboratory of Precision and Intelligent Chemistry, Hefei National Research Center for Physical Sciences at the Microscale, School of Chemistry and Materials Science, University of Science and Technology of China, Hefei, China
D
Daobin Liu
State Key Laboratory of Precision and Intelligent Chemistry, Hefei National Research Center for Physical Sciences at the Microscale, School of Chemistry and Materials Science, University of Science and Technology of China, Hefei, China
Y
Yuebo Liu
Center for Scientific Intelligence Innovation, Hefei, China; School of Chemistry and Materials Science, University of Science and Technology of China, Hefei, China
J
Jie Li
Center for Scientific Intelligence Innovation, Hefei, China; School of Chemistry and Materials Science, University of Science and Technology of China, Hefei, China
X
Xiaohui Li
Center for Scientific Intelligence Innovation, Hefei, China; School of Chemistry and Materials Science, University of Science and Technology of China, Hefei, China
Linjiang Chen
Linjiang Chen
USTC & Bham
Yi Luo
Yi Luo
Assistant Professor/Member, Moffitt Cancer Center
Machine learningsystem informaticshealth outcomesdecision supportcredible model
Jun Jiang
Jun Jiang
University of Science and Technology of China
Theoretical ChemistryPhysical ChemistryPhotocatalysis/CatalysisMaterial Design