🤖 AI Summary
This study addresses the challenge of reaction mechanism prediction during cross-chemical-domain transfer, where molecular unfamiliarity and data scarcity hinder performance. We propose MechaVLM, a framework that pioneers an open-book prediction paradigm by directly leveraging visual literature images rather than symbolic parsing. By integrating transferable visual representations with external mechanistic knowledge through multi-scale chemical grounding, cross-rendering contrastive learning, and precedent retrieval, MechaVLM recursively generates electron edits to construct complete mechanisms. Additionally, we introduce the MechBench benchmark to enable zero-shot mechanism generation without retraining. Experimental results demonstrate that our framework improves Top-1 accuracy by over 12% on the FlowER-to-ReactMech transfer task, significantly enhances out-of-distribution reaction prediction, and exhibits strong generalization capabilities in tasks such as atom mapping.
📝 Abstract
Reaction mechanisms describe the step-by-step transformations underlying chemical reactions and are central to reaction analysis and synthesis. Learning-based models have achieved strong performance on established mechanism-prediction benchmarks, but transferring them to unseen chemistry remains challenging. Such transfer is difficult because familiar mechanisms must be applied to unfamiliar molecular structures, and some target mechanisms may be poorly covered by the training data. To address these challenges, we introduce MechaVLM, a visual framework that combines transferable chemical representations with external mechanistic knowledge. It learns reusable visual features through multiscale chemical grounding and cross-rendering contrastive learning. For open-book prediction, MechaVLM retrieves a fixed set of precedents from 70,384 literature mechanism figures and re-reads relevant visual evidence as the molecular state evolves, directly using the figures without symbolic mechanism parsing. An atom-indexed language decoder then recursively generates executable electron edits to construct the complete mechanism. We further introduce MechBench, a challenging literature-derived benchmark with 2,184 mechanisms and 9,146 elementary steps. Across cross-dataset and literature-derived benchmarks, MechaVLM establishes strong zero-shot mechanism prediction. Its closed-book model alone improves Step/Pathway Top-1 by 12.50/13.93 percentage points on FlowER-to-ReactMech transfer, while external visual precedents unlock further gains on challenging OOD reactions. The learned representation also generalizes beyond mechanism prediction to atom mapping and reaction center prediction.