🤖 AI Summary
Traditional reaction yield prediction relies on one-dimensional quantum descriptors that lack spatial information, making it difficult to accurately model molecular steric effects. This work proposes a bimodal visual cross-attention architecture that, for the first time, leverages general-purpose computer vision backbone networks to represent two-dimensional molecular topology images. These visual representations are fused with tabular physical organic data, and a cross-attention mechanism dynamically learns chemical hierarchies to focus on critical steric bottlenecks, while residual connections preserve non-spatial electronic parameters. Evaluated on a test set, the method achieves an RMSE of 5.27%, significantly outperforming baseline models based solely on quantum descriptors, and offers both higher predictive accuracy and enhanced interpretability.
📝 Abstract
Traditional reaction yield prediction is constrained by 1D quantum descriptors that lack explicit spatial information. To address this gap, a dual-modal Vision Cross-Attention architecture is proposed, fusing tabular physical-organic data with 2D molecular topologies. Notably, it is demonstrated that a generic computer vision backbone processing simple 2D skeletal structures independently outperforms purely quantum-based baselines. By synergizing both modalities, superior predictive accuracy compared to traditional methodologies is achieved by the optimal cross-attention framework (Test RMSE = 5.27%). Through mechanistic probing, active, descriptor-guided spatial querying is observed, effectively offloading macroscopic steric identification to the visual pathway. Furthermore, a dynamic chemical hierarchy is learned by the network to heavily prioritize critical steric bottlenecks, such as the aryl halide. Concurrently, residual skip connections are utilized to protect non-spatial electronic parameters from destructive attenuation during fusion. Collectively, a scalable and highly interpretable blueprint is provided for augmenting physical chemistry with deep visual learning.