Cutting Through the Noise: On-the-fly Outlier Detection for Robust Training of Machine Learning Interatomic Potentials

📅 2026-02-09
📈 Citations: 0
Influential: 0
📄 PDF

career value

231K/year
🤖 AI Summary
This work addresses the performance degradation of machine learning interatomic potentials caused by noise from unconverged or inconsistent electronic structure calculations in training data. Existing denoising approaches rely on manual curation or iterative retraining, which are computationally expensive. To overcome this limitation, the authors propose an unsupervised online denoising method that dynamically tracks the loss distribution during a single training run using exponential moving averages, enabling real-time detection and automatic down-weighting of anomalous samples without requiring additional reference calculations or iterative retraining. The method successfully recovers accurate diffusion coefficients from unconverged liquid water data and reduces energy prediction errors by a factor of three on the SPICE dataset of organic molecules, demonstrating high efficiency, scalability, and robustness.

Technology Category

Application Category

📝 Abstract
The accuracy of machine learning interatomic potentials suffers from reference data that contains numerical noise. Often originating from unconverged or inconsistent electronic-structure calculations, this noise is challenging to identify. Existing mitigation strategies such as manual filtering or iterative refinement of outliers, require either substantial expert effort or multiple expensive retraining cycles, making them difficult to scale to large datasets. Here, we introduce an on-the-fly outlier detection scheme that automatically down-weights noisy samples, without requiring additional reference calculations. By tracking the loss distribution via an exponential moving average, this unsupervised method identifies outliers throughout a single training run. We show that this approach prevents overfitting and matches the performance of iterative refinement baselines with significantly reduced overhead. The method's effectiveness is demonstrated by recovering accurate physical observables for liquid water from unconverged reference data, including diffusion coefficients. Furthermore, we validate its scalability by training a foundation model for organic chemistry on the SPICE dataset, where it reduces energy errors by a factor of three. This framework provides a simple, automated solution for training robust models on imperfect datasets across dataset sizes.
Problem

Research questions and friction points this paper is trying to address.

outlier detection
machine learning interatomic potentials
numerical noise
robust training
data quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

on-the-fly outlier detection
machine learning interatomic potentials
unsupervised anomaly detection
robust training
exponential moving average
🔎 Similar Papers
T
Terry C. W. Lam
Cavendish Laboratory, Department of Physics, University of Cambridge, Cambridge, CB3 0HE, United Kingdom; Lennard-Jones Centre, University of Cambridge, Trinity Lane, Cambridge, CB2 1TN, United Kingdom
N
Niamh O'Neill
Yusuf Hamied Department of Chemistry, University of Cambridge, Lensfield Road, Cambridge, CB2 1EW, United Kingdom; Cavendish Laboratory, Department of Physics, University of Cambridge, Cambridge, CB3 0HE, United Kingdom; Lennard-Jones Centre, University of Cambridge, Trinity Lane, Cambridge, CB2 1TN, United Kingdom
C
Christoph Schran
Cavendish Laboratory, Department of Physics, University of Cambridge, Cambridge, CB3 0HE, United Kingdom; Lennard-Jones Centre, University of Cambridge, Trinity Lane, Cambridge, CB2 1TN, United Kingdom
L
Lars L. Schaaf
Cavendish Laboratory, Department of Physics, University of Cambridge, Cambridge, CB3 0HE, United Kingdom; Lennard-Jones Centre, University of Cambridge, Trinity Lane, Cambridge, CB2 1TN, United Kingdom