🤖 AI Summary
This study addresses the lack of robustness of Shapley values in tree-based models to minor data perturbations, which often leads to biased explanations. To mitigate this issue, we propose the first interval-valued SHAP framework for tree models based on the Imprecise Dirichlet Model (IDM). By incorporating unlabeled instances to quantify interval attributions and combining pessimistic and average principles to handle incomplete data, our approach effectively captures attribution uncertainty. Furthermore, we design an efficient algorithm that extends this framework to Banzhaf value computation. The proposed method significantly enhances the stability of feature attribution and the robustness of model explanations while effectively eliminating biases introduced by uninformative features. Overall, this work provides a reliable interval-valued interpretability solution for tree-based models.
📝 Abstract
Shapley values are among the most popular feature-attribution explanations. Efficient approaches for computing/estimating Shapley values for tree-based models, which are state-of-the-art for tabular data sets, have been developed. However, it is known that Shapley values can be (highly) unrobust due to small and realistic changes. In this paper, we propose an imprecise Dirichlet model (IDM) based method to analyze the robustness of Shapley values in decision trees and random forests. Technically, it is done by quantifying and analyzing the interval-valued Shapley values when a few unannotated instances are randomly introduced to the leaves of the trees. The interval-valued Shapley values can be defined following common principles in handling incomplete data: the pessimistic and averaging principles. We derive various theoretical results that lead to efficient computation of the interval-valued Shapley values. We also show that the proposed method can be straightforwardly generalized to the case of Banzhaf values. We then present various case studies and experiments to illustrate the behaviour of the proposed interval-valued Shapley values and their applications in debiasing uninformative features.