🤖 AI Summary
This work addresses the challenge of effectively integrating electronic descriptors with the three-dimensional geometric structure of active sites in transition-metal-catalyzed reactions. To this end, we propose ChemFusion, a multimodal neural network that jointly encodes electronic features and explicit 3D atomic coordinates. Built upon an unpooled molecular point cloud representation, ChemFusion employs a cross-attention mechanism to enable end-to-end modeling of electronic states and spatial configurations, while attention weights automatically highlight critical steric hindrance regions, enhancing the model’s physical interpretability. Evaluated across multiple cross-coupling reaction datasets, ChemFusion significantly outperforms unimodal approaches, achieving higher accuracy in yield prediction and revealing steric effects that govern reaction selectivity.
📝 Abstract
Forecasting the outcomes of transition-metal-catalyzed reactions is notoriously complex due to the interplay of diverse physical and chemical variables. A persistent computational bottleneck has been effectively merging broad electronic descriptors with the localized, three-dimensional geometry of the reactive site. To bridge this representation gap, we present ChemFusion, a hybrid neural network that fuses conventional electronic features with explicit 3D atomic coordinates. Using a cross-attention mechanism, the model enables global electronic states to dynamically attend to specific spatial constraints within un-pooled molecular point clouds. When benchmarked against a diverse library of cross-couplings, this approach delivers exceptional predictive performance, decisively surpassing traditional single-modality frameworks. Importantly, extracting the attention matrices reveals that the architecture autonomously learns to identify and penalize restrictive steric hindrances. This provides a physically grounded interpretability, demonstrating that spatially aware networks can navigate complex reaction sterics that standard statistical models typically miss.