๐ค AI Summary
This study addresses the limited cross-disciplinary generalization of existing atomistic generative models and their insufficient integration of organic and inorganic data by proposing a universal atomistic generative model pretrained on five million structures. Methodologically, it introduces the first unified organic-inorganic pretraining framework to enable cross-domain transfer learning, and employs a multi-scale Transformer architecture combined with conditional flow matching and force-field conditioning mechanisms to support multi-scale generation tasks. The proposed model significantly improves molecular distribution fidelity and increases the success rate of protein backbone design from 67.8% to 74.8% under few-shot conditions. Ultimately, this work provides a unified and efficient solution for cross-domain atomistic generation.
๐ Abstract
Unified atomistic modeling has the potential to accelerate discovery in chemistry, materials science, and biology by bridging data-rich chemical domains and data-scarce biological contexts. However, existing generative approaches to atomistic modeling remain highly specialized to scientific disciplines (chemistry vs. biology) or do not leverage both high-volume organic (molecule) and inorganic (material) data for general-purpose pretraining. To this end, we introduce Zatom-2, an atomistic generative model pretrained on approximately five million structures from the OMol25 and OMat24 electronic structure datasets. Zatom-2 features a multiscale Transformer architecture coupled with conditional flow matching that supports force conditioning and foundational pretraining tasks such as generation, structure prediction, and prediction of molecular and material energies and forces. Empirically, Zatom-2 achieves better molecular distribution fidelity than Zatom-1 and achieves strong performance on existing molecule and material generation benchmarks. Zatom-2 demonstrates the ability to control sample generation across low- and high-force regimes, and enhances protein generation in a low-data setting through joint generative-predictive pretraining and transfer learning, increasing protein backbone designability in a length extrapolation setting from 67.8% without pretraining to 74.8% after finetuning on 2,000 protein domains.