🤖 AI Summary
Conventional machine-learned interatomic potentials (MLIPs) employ scalar regression to predict energies and forces, achieving high efficiency but lacking principled quantification of epistemic uncertainty.
Method: This work introduces the first classification-based MLIP framework, mapping continuous energy and force targets onto histogram-binned distributional labels and jointly optimizing probabilistic outputs via cross-entropy loss. Uncertainty is intrinsically quantified through the entropy of predicted distributions, enabling direct, interpretable confidence estimation.
Contribution/Results: On DFT-derived benchmark datasets, the method achieves absolute prediction errors comparable to state-of-the-art regression-based MLIPs while providing calibrated, distribution-level uncertainty estimates. It overcomes the long-standing limitation of uncertainty-unaware modeling in MLIPs and establishes a risk-aware paradigm for molecular simulation—enabling reliability assessment, active learning, and safe decision-making in computational chemistry and materials science.
📝 Abstract
Density Functional Theory (DFT) is a widely used computational method for estimating the energy and behavior of molecules. Machine Learning Interatomic Potentials (MLIPs) are models trained to approximate DFT-level energies and forces at dramatically lower computational cost. Many modern MLIPs rely on a scalar regression formulation; given information about a molecule, they predict a single energy value and corresponding forces while minimizing absolute error with DFT's calculations. In this work, we explore a multi-class classification formulation that predicts a categorical distribution over energy/force values, providing richer supervision through multiple targets. Most importantly, this approach offers a principled way to quantify model uncertainty.
In particular, our method predicts a histogram of the energy/force distribution, converts scalar targets into histograms, and trains the model using cross-entropy loss. Our results demonstrate that this categorical formulation can achieve absolute error performance comparable to regression baselines. Furthermore, this representation enables the quantification of epistemic uncertainty through the entropy of the predicted distribution, offering a measure of model confidence absent in scalar regression approaches.