🤖 AI Summary
This study addresses the frequent failure of machine learning models in screening cathode materials for sodium-ion batteries, which often stems from reliance on computationally derived reference voltages—such as those from PBE+U—that exhibit systematic errors. Employing a preregistered experimental validation protocol, we evaluate a graph neural network model against experimental measurements across six known materials and quantify the discrepancy between computed and empirical voltages. Our results reveal that Materials Project’s PBE+U voltages are systematically underestimated by approximately 0.54 V, constituting the dominant source of model error (with a holdout-set MAE of 0.67 V and a 95% confidence upper bound of 1.09 V), and this bias strongly correlates with the voltage magnitude. To mitigate such issues, we propose a calibration and auditing framework tailored to DFT-based data ledgers, establishing a more reliable benchmark for computational materials screening.
📝 Abstract
Machine-learning screens for battery materials are trained and judged almost entirely against computed reference voltages, and those references carry their own systematic errors. We report a case in which this matters quantitatively: our own screening stack (a graph-network voltage screen, a prior-art triage layer, and a local PBE+U bench) fails pre-registered validation against experiment-anchored literature values. Verdict thresholds, failure modes, and the primary metric were committed before analysis. On an operator-audited set of known Na-ion cathodes (n = 6 after one documented exclusion; verdict unchanged at n = 7), the raw held-out mean absolute error was 0.67 V, the pre-registered conservative metric, the upper 95% confidence bound of the cross-validated bias-corrected error, was 1.09 V, and the residual was strongly voltage-dependent (r = -0.94), so no additive calibration is valid. On the two compounds where prediction, database reference, and experiment could all be compared, the Materials Project PBE+U reference sat about 0.54 V below measurement: the reference, not the model, dominated the error. A prior-art screen found at least 70% of the targeted Na substitution space already published. We retire the screen, bound what "verified" means for our DFT ledger, and pre-register a calibration audit of it against four benchmark Li couples.