🤖 AI Summary
This study addresses the inconsistent feature attribution and lack of geometry-aware interpretability caused by geometric operations in hyperbolic neural networks. We propose the LRP-radial-all rule, which leverages the Poincaré–Lorentz logarithmic map and inter-layer relevance propagation to handle the geometric scaling and signal branching within radial modules. By introducing geometric representation invariance (GRI) and a zero-curvature consistency criterion, our method overcomes the failure of traditional conservative propagation under equivalent computations, thereby ensuring relevance conservation and equivalence decomposition invariance. Experiments demonstrate that the proposed approach achieves high attribution fidelity on MNIST, sEEG, and CIFAR-10, with computational efficiency significantly surpassing Integrated Gradients and existing baseline methods.
📝 Abstract
Hyperbolic neural networks introduce geometric operations that require explicit treatment in relevance propagation. Equivalent geometric realizations can produce different feature attributions, even when local relevance is conserved. We study this problem through Geometric Representation Invariance (GRI), a specialization of Implementation Invariance, and zero-curvature consistency, which requires identity relevance propagation when a geometric module approaches the identity. We propose LRP-radial-all for origin-centered radial modules, treating geometric scaling as modulation and assigning relevance entirely to the signal branch. The rule conserves relevance, is invariant to equivalent radial factorizations, and satisfies zero-curvature consistency, yielding GRI for a specified Poincar\'e-Lorentz logarithmic-map construction. In contrast, a conservative LRP-half baseline can violate both consistency criteria. Experiments on hyperbolic MNIST, sEEG, and CIFAR-10 classifiers assess attribution fidelity, qualitative explanations, and runtime. LRP-radial-all achieves competitive attribution fidelity across datasets with runtime comparable to Gradient$\times$Input and substantially lower than Integrated Gradients. These findings motivate geometry-aware propagation rules that distinguish relevance conservation from consistency across equivalent computations.