๐ค AI Summary
This work addresses the limitation of existing graph Transformers, which rely on a single-token paradigm for graph-level representation and consequently fail to fully exploit the sequence modeling capacity of self-attention, often reducing to a weighted sum of node features. To overcome this, the authors propose a sequential graph tokenization paradigm that transforms node information into a sequence of tokens equipped with positional encodings. By stacking self-attention layers, the model captures complex dependencies among tokens, thereby unlocking the Transformerโs ability to model global structural information in graphs. This approach transcends the constraints of the conventional single-token framework and achieves state-of-the-art performance across multiple graph-level benchmark tasks. Ablation studies further confirm the effectiveness of each proposed component.
๐ Abstract
Transformers have demonstrated success in graph learning, particularly for node-level tasks. However, existing methods encounter an information bottleneck when generating graph-level representations. The prevalent single token paradigm fails to fully leverage the inherent strength of self-attention in encoding token sequences, and degenerates into a weighted sum of node signals. To address this issue, we design a novel serialized token paradigm to encapsulate global signals more effectively. Specifically, a graph serialization method is proposed to aggregate node signals into serialized graph tokens, with positional encoding being automatically involved. Then, stacked self-attention layers are applied to encode this token sequence and capture its internal dependencies. Our method can yield more expressive graph representations by modeling complex interactions among multiple graph tokens. Experimental results show that our method achieves state-of-the-art results on several graph-level benchmarks. Ablation studies verify the effectiveness of the proposed modules.