🤖 AI Summary
This paper addresses the lack of statistical reliability in global feature importance assessment for black-box models. We propose φ-test, the first method that integrates Shapley-value-guided feature selection with selective inference to enable interpretable and statistically verifiable global feature selection and significance testing. φ-test constructs a linear surrogate model and outputs a global importance table containing Shapley values, regression coefficients, post-selection p-values, and confidence intervals. In regression tasks involving tree-based and neural network models, the selected features—typically few in number—retain over 90% of the original model’s predictive performance. Moreover, the selected feature sets exhibit high stability across different base models and bootstrap resamples. By bridging SHAP-based explanation with classical statistical inference, φ-test establishes a new paradigm for explainable AI that jointly ensures interpretability and statistical rigor.
📝 Abstract
We propose $φ$-test, a global feature-selection and significance procedure for black-box predictors that combines Shapley attributions with selective inference. Given a trained model and an evaluation dataset, $φ$-test performs SHAP-guided screening and fits a linear surrogate on the screened features via a selection rule with a tractable selective-inference form. For each retained feature, it outputs a Shapley-based global score, a surrogate coefficient, and post-selection $p$-values and confidence intervals in a global feature-importance table. Experiments on real tabular regression tasks with tree-based and neural backbones suggest that $φ$-test can retain much of the predictive ability of the original model while using only a few features and producing feature sets that remain fairly stable across resamples and backbone classes. In these settings, $φ$-test acts as a practical global explanation layer linking Shapley-based importance summaries with classical statistical inference.