🤖 AI Summary
In materials informatics, nonlinear models suffer from poor interpretability, feature redundancy, and strong inter-feature correlations—leading to spurious feature importance estimates. To address this, we propose an XML-guided feature pruning framework that integrates SVR modeling with SHAP and permutation importance analysis, systematically identifying and removing redundant and spuriously important features on a GW-scale bandgap dataset; crucially, it emphasizes correlation-aware preprocessing for reliable attribution. The resulting lightweight 5-feature model maintains in-domain accuracy (MAE ≈ 0.15 eV) while reducing out-of-domain generalization error by over 30%—significantly outperforming the full-feature baseline. Our core contribution lies in uncovering the mechanistic interference of strong correlations on feature importance estimation and establishing a bandgap prediction paradigm that simultaneously achieves high accuracy, strong interpretability, and robust extrapolation capability.
📝 Abstract
In the rapidly advancing field of materials informatics, nonlinear machine learning models have demonstrated exceptional predictive capabilities for material properties. However, their black-box nature limits interpretability, and they may incorporate features that do not contribute to, or even deteriorate, model performance. This study employs explainable ML (XML) techniques, including permutation feature importance and the SHapley Additive exPlanation, applied to a pristine support vector regression model designed to predict band gaps at the GW level using 18 input features. Guided by XML-derived individual feature importance, a simple framework is proposed to construct reduced-feature predictive models. Model evaluations indicate that an XML-guided compact model, consisting of the top five features, achieves comparable accuracy to the pristine model on in-domain datasets while demonstrating superior generalization with lower prediction errors on out-of-domain data. Additionally, the study underscores the necessity for eliminating strongly correlated features to prevent misinterpretation and overestimation of feature importance before applying XML. This study highlights XML's effectiveness in developing simplified yet highly accurate machine learning models by clarifying feature roles.