🤖 AI Summary
This study addresses the long-standing lack of theoretical foundations for attributing and interpreting topological features, specifically persistent homology classes, in machine learning models. To this end, it proposes LandscapeSHAP, which introduces Shapley values from cooperative game theory into topological data analysis for the first time, providing a fair credit assignment scheme for points in persistence diagrams. For linear models, the authors derive closed-form solutions and design an efficient Monte Carlo sampling algorithm. Furthermore, this work establishes the axiomatic uniqueness of the proposed attribution method, enabling both exact and approximate attributions. It also theoretically guarantees convergence rates and stability. Collectively, these contributions fill a critical gap in the field of topological feature attribution by offering a rigorous, interpretable framework for understanding how topological structures influence model predictions.
📝 Abstract
Shapley values, a solution concept from cooperative game theory, have recently become a standard tool for feature credit allocation in machine learning. They provide an axiomatically justified method to fairly distribute a model's prediction among the data features. Shapley values have not yet been applied to explain machine learning models trained on features from topological data analysis. We develop what we believe is the first such approach, focusing on the persistence landscape featurization of persistence diagrams. Because each landscape coordinate is a rank statistic, crediting a model's prediction back to individual persistent homology classes (persistence diagram points) is nontrivial. We introduce LandscapeSHAP, a method for fair credit allocation to persistence diagram points based on a model's prediction. For linear models on persistence landscapes, LandscapeSHAP has a closed form expression that gives the exact Shapley value of every persistence diagram point. In particular, there is no coalition sampling required. We further prove that the four Shapley "fairness" axioms uniquely characterize this credit allocation for any model, not only linear ones. For a general nonlinear model, this unique value can only be calculated exactly from its defining coalition averaging formula, which requires considering all $2^N$ many coalitions, where $N$ is the number of points in the persistence diagram. This is computationally intractable for persistence diagrams of realistic size. We complement the exact linear model result with an efficient Monte Carlo sampling of persistence diagram coalitions. We give convergence rates in terms of number of samples needed to approximate to a desired degree of accuracy. We also prove stability results for the LandscapeSHAP credit allocation, for any model.