🤖 AI Summary
Existing machine-learned interatomic potentials (MLIPs) for heterogeneous catalysis modeling are constrained by limited training data and struggle to simultaneously satisfy spin-polarization awareness and high-fidelity accuracy requirements. To address this, we introduce AQCat25—a large-scale dataset comprising 13.5 million DFT calculations—and propose a meta-data conditional modulation mechanism based on Feature-wise Linear Modulation (FiLM), which explicitly encodes system-specific attributes such as spin state and exchange-correlation functional type. This design mitigates knowledge conflicts and catastrophic forgetting inherent in multi-fidelity and multi-physics modeling. By jointly training with the OC20 dataset, our model preserves strong generalization across diverse catalytic systems while significantly improving performance on spin-sensitive and high-accuracy tasks—particularly transition-state energy barrier prediction. To our knowledge, this work presents the first general-purpose catalytic potential that concurrently achieves spin-awareness, high predictive fidelity, and robust cross-dataset transferability.
📝 Abstract
Large-scale datasets have enabled highly accurate machine learning interatomic potentials (MLIPs) for general-purpose heterogeneous catalysis modeling. There are, however, some limitations in what can be treated with these potentials because of gaps in the underlying training data. To extend these capabilities, we introduce AQCat25, a complementary dataset of 13.5 million density functional theory (DFT) single point calculations designed to improve the treatment of systems where spin polarization and/or higher fidelity are critical. We also investigate methodologies for integrating new datasets, such as AQCat25, with the broader Open Catalyst 2020 (OC20) dataset to create spin-aware models without sacrificing generalizability. We find that directly tuning a general model on AQCat25 leads to catastrophic forgetting of the original dataset's knowledge. Conversely, joint training strategies prove effective for improving accuracy on the new data without sacrificing general performance. This joint approach introduces a challenge, as the model must learn from a dataset containing both mixed-fidelity calculations and mixed-physics (spin-polarized vs. unpolarized). We show that explicitly conditioning the model on this system-specific metadata, for example by using Feature-wise Linear Modulation (FiLM), successfully addresses this challenge and further enhances model accuracy. Ultimately, our work establishes an effective protocol for bridging DFT fidelity domains to advance the predictive power of foundational models in catalysis.