Measuring Explainer Stability via Attribution Separability

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing attribution methods suffer from instability due to inherent randomness and lack a quantitative measure for the reliability of feature rankings. This work proposes a distribution-based framework to evaluate attribution stability, introducing the novel concept of “attribution separability.” By modeling the distribution of attribution vectors and quantifying their separability, the method identifies the highest-ranked index up to which feature ordering remains reliable. The proposed framework establishes a new quantitative criterion for interpreter stability and enables robust cross-dataset comparisons of ranking reliability across different attribution methods. Experimental results demonstrate that the framework effectively assesses and compares the stability of diverse attribution techniques, offering a principled basis for selecting interpreters in practice.
📝 Abstract
Attribution methods (AMs) assign an importance score to each feature and are widely adopted to explain black-box models. However, most methods can produce variable attribution scores due to stochastic components in their definition. In this paper, we propose a distribution-based framework to capture the stability of attribution scores. In particular, our approach allows to understand the degree of separability in the ranked attribution vector and obtain the largest index for which a feature ranking remains reliable. We further extend this framework to compare AMs based on the robustness of their rankings across a dataset. Through experiments, we demonstrate how to apply our method to evaluate explainer stability. Overall, our approach provides a complementary criterion for evaluating the stability of AMs.
Problem

Research questions and friction points this paper is trying to address.

attribution methods
explainer stability
attribution separability
feature ranking reliability
black-box models
Innovation

Methods, ideas, or system contributions that make the work stand out.

attribution stability
attribution separability
feature ranking robustness
distribution-based evaluation
explainable AI