RxnOptBench: Benchmarking LLMs for Reaction-Condition Optimization in Organic Methodology

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of benchmarks evaluating large language models (LLMs) in optimizing reaction conditions using real experimental data. To this end, we introduce RxnOptBench, constructed from wet-lab data in 2025 organic methodology publications, to assess model capabilities in maximizing yield and stereoselectivity through catalyst and solvent selection. Key innovations include the first continuous scoring mechanism based on yield and ee/dr/rr values, a precedent/no-precedent control design to decouple memorization from reasoning, and an evaluation framework integrating literature mining with multidimensional metrics. Evaluations across nine state-of-the-art models reveal suboptimal performance: specialized chemistry models perform near random baselines, while open-source models substantially narrow the gap with proprietary counterparts.
📝 Abstract
Chemical reaction-condition optimization -- choosing the catalyst, ligand, solvent, reagent, temperature, time, and atmosphere that jointly maximize yield and stereoselectivity -- is a central, judgement-laden subtask of organic methodology research that large language models are increasingly expected to support. Yet existing chemistry benchmarks evaluate reaction-class labelling, retrosynthesis, or SMILES manipulation, and do not ask models to read a real condition-screening table and pick the best set. We introduce RxnOptBench, a benchmark whose every option and precedent is a real wet-lab entry mined from the optimization tables of organic-methodology papers published in 2025, graded by a continuous relative score derived from a declared headline utility that combines reported yield with enantiomeric excess (ee), diastereomeric ratio (dr), and regioisomeric ratio (rr), and equipped with a paired precedents-vs-no-precedents design that isolates in-context use of literature evidence from parametric memorization. Across nine frontier LLMs and three Chemistry LLMs, even the best models leave substantial headroom: chemistry-specialized models fall to the random-baseline floor on multi-axis selection, while open-weight models have closed most of the gap to proprietary frontier models. We release the final human-reviewed benchmark test set and evaluation code.
Problem

Research questions and friction points this paper is trying to address.

reaction-condition optimization
large language models
benchmark
organic methodology
stereoselectivity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reaction-condition optimization
Benchmark
Large language models
Organic methodology
In-context learning
💼 Related Jobs
No related jobs found.
L
Lingli Ge
Shanghai Jiao Tong University, Shanghai, China
Yubin Wang
Yubin Wang
Shanghai AI Lab
deep learningcomputer visionperson re-id
J
Junyuan Gao
Shanghai Artificial Intelligence Laboratory, Shanghai, China
Jiahe Song
Jiahe Song
SJTU&Shanghai AI Lab, PhD Candidate
LLMVLM
J
Jiaxing Sun
Shanghai Artificial Intelligence Laboratory, Shanghai, China
B
Boyu Zhu
Shanghai Artificial Intelligence Laboratory, Shanghai, China
Haote Yang
Haote Yang
PJLab
CVLLMMLLMAI4S
Jingchao Wang
Jingchao Wang
East China Normal University
AI
L
Lixin Ma
Shanghai Artificial Intelligence Laboratory, Shanghai, China
Jiang Wu
Jiang Wu
Shanghai Artificial Intelligence Laboratory
large language modelvision language model
Yuqiang Li
Yuqiang Li
Central South University
Internal Combustion EngineCombustionEmissionsMechansim
Conghui He
Conghui He
Shanghai AI Laboratory
Data-centric AILLMDocument Intelligence