From Regression to Classification: Exploring the Benefits of Categorical Representations of Energy in MLIPs

📅 2025-11-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Conventional machine-learned interatomic potentials (MLIPs) employ scalar regression to predict energies and forces, achieving high efficiency but lacking principled quantification of epistemic uncertainty. Method: This work introduces the first classification-based MLIP framework, mapping continuous energy and force targets onto histogram-binned distributional labels and jointly optimizing probabilistic outputs via cross-entropy loss. Uncertainty is intrinsically quantified through the entropy of predicted distributions, enabling direct, interpretable confidence estimation. Contribution/Results: On DFT-derived benchmark datasets, the method achieves absolute prediction errors comparable to state-of-the-art regression-based MLIPs while providing calibrated, distribution-level uncertainty estimates. It overcomes the long-standing limitation of uncertainty-unaware modeling in MLIPs and establishes a risk-aware paradigm for molecular simulation—enabling reliability assessment, active learning, and safe decision-making in computational chemistry and materials science.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationReasoning under Uncertainty: Uncertainty RepresentationsMultiagent Systems: Multiagent Systems under Uncertainty

Application Category

Responsible Web: Machine-in-the-loop, human agency and autonomySemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
Density Functional Theory (DFT) is a widely used computational method for estimating the energy and behavior of molecules. Machine Learning Interatomic Potentials (MLIPs) are models trained to approximate DFT-level energies and forces at dramatically lower computational cost. Many modern MLIPs rely on a scalar regression formulation; given information about a molecule, they predict a single energy value and corresponding forces while minimizing absolute error with DFT's calculations. In this work, we explore a multi-class classification formulation that predicts a categorical distribution over energy/force values, providing richer supervision through multiple targets. Most importantly, this approach offers a principled way to quantify model uncertainty. In particular, our method predicts a histogram of the energy/force distribution, converts scalar targets into histograms, and trains the model using cross-entropy loss. Our results demonstrate that this categorical formulation can achieve absolute error performance comparable to regression baselines. Furthermore, this representation enables the quantification of epistemic uncertainty through the entropy of the predicted distribution, offering a measure of model confidence absent in scalar regression approaches.
Problem

Research questions and friction points this paper is trying to address.

Exploring categorical energy representation in MLIPs
Quantifying model uncertainty via classification formulation
Achieving comparable accuracy to regression with richer supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Categorical distribution predicts energy/force histograms
Cross-entropy loss trains model using multiple targets
Quantifies uncertainty via entropy of predicted distribution
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.