🤖 AI Summary
This study addresses the limitation of existing GNN explanation methods that optimize fidelity while neglecting interpretability and stability. We propose a controllable explanation framework based on multi-metric preference selection, which jointly optimizes fidelity, interpretability, and stability, achieving controllable optimization under multi-objective trade-offs for the first time. Specifically, exposure weights are employed as a control mechanism, and a domain-prior motif library is introduced to guide subgraph pattern matching, significantly enhancing both explanation quality and search efficiency. Experimental results demonstrate that our method yields explanations with higher fidelity across multiple benchmark datasets, achieves superior inference speed under limited computational budgets, and effectively reveals the intrinsic trade-off relationships among various evaluation metrics.
📝 Abstract
Mechanisms for generating GNN explanations are crucial for building trust and mitigating biases in Graph Neural Networks (GNNs), especially in high-stakes scenarios. Most current methods optimize only for fidelity under the sparsity constraint. However, this discounts the need for interpretable explanations (those that consist of familiar motif patterns) and stable explanations (those that remain unchanged under structural perturbations). We propose a novel approach that optimizes GNN explanations across these metrics, exposing their relative weighing as a control. Experiments on various real-world datasets, including MUTAG, BA-2Motif, BAMultiShapes, and PROTEINS, suggest that our method produces higher-fidelity explanations than a state-of-the-art baseline on MUTAG and PROTEINS across all evaluated budgets, and on BA-2Motif at larger budgets, while being faster in the regime of small explanation budgets. We also explore how, given an input motif library containing standard motifs for the corresponding domain, the method can be used to determine the relative importance of those motifs in generating the explanations, and how this information can be used to further improve the quality of the output explanations. We also examine the relationship between different metrics through their induced tradeoff surface, and explore its dependence on the nature of the motif library.