🤖 AI Summary
This study addresses the lack of benchmarks evaluating large language models (LLMs) in optimizing reaction conditions using real experimental data. To this end, we introduce RxnOptBench, constructed from wet-lab data in 2025 organic methodology publications, to assess model capabilities in maximizing yield and stereoselectivity through catalyst and solvent selection. Key innovations include the first continuous scoring mechanism based on yield and ee/dr/rr values, a precedent/no-precedent control design to decouple memorization from reasoning, and an evaluation framework integrating literature mining with multidimensional metrics. Evaluations across nine state-of-the-art models reveal suboptimal performance: specialized chemistry models perform near random baselines, while open-source models substantially narrow the gap with proprietary counterparts.
📝 Abstract
Chemical reaction-condition optimization -- choosing the catalyst, ligand, solvent, reagent, temperature, time, and atmosphere that jointly maximize yield and stereoselectivity -- is a central, judgement-laden subtask of organic methodology research that large language models are increasingly expected to support. Yet existing chemistry benchmarks evaluate reaction-class labelling, retrosynthesis, or SMILES manipulation, and do not ask models to read a real condition-screening table and pick the best set. We introduce RxnOptBench, a benchmark whose every option and precedent is a real wet-lab entry mined from the optimization tables of organic-methodology papers published in 2025, graded by a continuous relative score derived from a declared headline utility that combines reported yield with enantiomeric excess (ee), diastereomeric ratio (dr), and regioisomeric ratio (rr), and equipped with a paired precedents-vs-no-precedents design that isolates in-context use of literature evidence from parametric memorization. Across nine frontier LLMs and three Chemistry LLMs, even the best models leave substantial headroom: chemistry-specialized models fall to the random-baseline floor on multi-axis selection, while open-weight models have closed most of the gap to proprietary frontier models. We release the final human-reviewed benchmark test set and evaluation code.