Beyond Trained Models: Compiling GNNs for a Sound Explainer Benchmark

πŸ“… 2026-10-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the lack of rigor in existing Graph Neural Network (GNN) explainer evaluations, which rely on unreliable presuppositions. To overcome this limitation, this work proposes a paradigm shift from black-box training to white-box compilation by pioneering the direct compilation of graded modal logic formulas into GNN weights. This approach constructs models with known behaviors to establish precise ground-truth explanations, upon which the GracrBench benchmark is developed. Evaluations using this benchmark reveal that most explainers lack robustness to indirect influences and varying implementation strategies. Ultimately, this research establishes a new standard for fine-grained diagnostics in evaluating GNN interpretability.
πŸ“ Abstract
Explainers for Graph Neural Networks (GNNs) are commonly evaluated by their plausibility, i.e., how well their explanations recover a predefined ground truth, such as a motif planted in the data. This protocol implicitly assumes that a GNN trained on such data relies on the intended motif. Although prior work has questioned this assumption, plausibility remains widespread. First, we show that the assumption is violated on several widely used benchmarks, where, e.g., degree statistics alone suffice to solve the task. Then, we remove this confounder by replacing training with compilation. We achieve this by introducing $\mathsf{Gracr}$, the first compiler translating graded modal logic formulas into GNN weights, yielding models that replicate the behaviour of the corresponding formulas. Since the behaviour of the model is now known by construction, we can define its ground truth explanation formally and compute it exactly. Building on this, we introduce $\mathsf{Gracr}\mathsf{Bench}$, a benchmark of compiled GNNs for the evaluation of explainers against this exact ground truth. Experiments on eleven explainers across six tasks show its effectiveness for fine-grained diagnostic evaluation: notably, we discover that most explainers are not robust to indirect influences or alternative implementations of the same formula. These results position $\mathsf{Gracr}\mathsf{Bench}$ as a novel, rigorous evaluation setting for graph post-hoc explainability.
Problem

Research questions and friction points this paper is trying to address.

Graph Neural Networks
Explainability
Evaluation Benchmark
Plausibility
Ground Truth
Innovation

Methods, ideas, or system contributions that make the work stand out.

Graph Neural Networks
Explainability Benchmark
Graded Modal Logic Compiler
Ground Truth Explanation
Post-hoc Explainability
πŸ”Ž Similar Papers
No similar papers found.