FREA: A Multi-Source Expert Benchmark for Reaction Feasibility Verification

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited cross-source generalization and misalignment with expert decision-making in AI-generated chemical reaction validators. We construct a multi-source reaction benchmark comprising 751 expert-annotated instances and systematically evaluate the verification performance of large language models and forward prediction models. By employing diverse negative sampling techniques for cross-analysis, we expose the generalization deficiencies of existing validators and release a corpus containing tens of millions of negative samples. Experimental results demonstrate that no single validator consistently outperforms others across all sources. Furthermore, while training on mixed negative samples improves overall AUROC, such gains fail to transfer effectively to model-proposed scenarios.
📝 Abstract
As generative models and AI agents propose chemical reactions at a scale beyond expert review, feasibility verifiers decide which proposals enter synthesis planning. But do their decisions agree with chemists across different kinds of candidates? We introduce FREA, a benchmark of 751 reactions labeled by expert chemists under an explicit feasibility criterion, drawn from retrosynthesis model proposals, zero-yield experimental records, edits by large language models (LLMs), and five negative candidate generation methods. Our evaluation finds that no verifier leads across all sources: LLMs given only the criterion are competitive with dedicated verifiers, while forward models perform best on retrosynthesis proposals but reject most feasible edits of recorded reactions at the evaluated operating points. Looking beyond aggregate scores, both forward models perform below chance when separating infeasible alternative disconnections from feasible generated candidates. To study whether negative supervision addresses these weaknesses, we also release a corpus of over 14 million recorded reactions and generated negative candidates. In matched training comparisons, adding a mixture of generated negatives to forward training raises mean AUROC across sources, but these gains do not extend to retrosynthesis proposals. Varying the generation method further shows that the largest gain on generated candidates coincides with worse proposal screening. These findings motivate evaluating verifiers against experts across sources and designing negatives for transfer to model proposals.
Problem

Research questions and friction points this paper is trying to address.

reaction feasibility verification
benchmark evaluation
generative models
retrosynthesis
expert annotation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reaction Feasibility Verification
Multi-Source Benchmark
Negative Supervision
Retrosynthesis
Forward Models
🔎 Similar Papers
No similar papers found.