🤖 AI Summary
This study aims to clarify the actual role of shared memory and global communication mechanisms in Graph Transformers for transductive learning on static graphs. Through theoretical analysis and empirical comparisons with message-passing neural networks, we demonstrate that under fixed graph settings, shared memory inherently degenerates into a constant, yielding performance comparable to directly optimized global communication. Furthermore, local models are shown capable of embedding equivalent global information. These findings establish that local architectures suffice to replace global communication mechanisms, thereby challenging prevailing evaluation paradigms and exposing limitations in current assessment methodologies for static graphs. Ultimately, this work offers a novel perspective for understanding the underlying mechanisms of Graph Transformers, suggesting that their purported advantages in transductive static graph settings may be largely attributable to factors other than global attention.
📝 Abstract
Scalable Graph Transformers are commonly trained and evaluated on static large graphs in a transductive setup. Many scalable Graph Transformer components can be formulated as a constant-size shared memory, similar to virtual nodes, providing compressed information about the whole graph. The counterpart of these models in language models and other domains is justified as the input changes, and this mechanism learns to compress some useful information about the input. In transductive learning on a single fixed graph, however, any shared memory can be seen as a constant at test time. This raises the question of what exactly this shared memory does in this static setup. We give preliminary evidence that optimizing a shared memory directly performs similarly to global communication methods, and so normal local message-passing models can embed similar information in their weights. Thus, these settings may be a poor fit for evaluating global communication in graph neural networks.