Attention Graphons: A Graph Limit Perspective on Graph Transformers

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether the attention matrices of Graph Transformers converge to a stable limit as the number of nodes grows. Grounded in dense graph limit theory, this work treats attention as samples from an underlying kernel function and introduces the concept of the "attention graphon." By employing cut-distance analysis, nonparametric estimation, and a block-averaging pipeline, it establishes assumption-free worst-case variance bounds alongside tighter regularity-based bounds. Empirical results confirm that attention converges to stable structures on specific datasets, validating the theoretical bounds and enabling cross-scale transferability. Ultimately, this research provides rigorous theoretical foundations for understanding the asymptotic behavior of large-scale Graph Transformers.
📝 Abstract
Graph Transformers produce, for each attention head, a dense $n\times n$ matrix of learned pairwise interactions. We ask a fundamental question: do these attention-induced graphs converge to a stable limit object as $n$ grows, or does the learned interaction pattern remain unstructured and size-dependent? We answer this using dense graph limit theory, treating each attention matrix as a finite sample from an underlying kernel---an \emph{attention graphon}---and studying concentration around this limit under the cut-distance. We derive a worst-case variance bound requiring no assumptions on the graphon, and a sharper regularity-aware bound based on nonparametric estimation theory. To operationalize the theory, we propose a canonicalize-then-block-average pipeline for estimating dataset-level attention graphons, and a variance-based diagnostic for testing whether attention admits a stable continuum description. Experiments across multiple graph benchmarks show that learned attention stabilizes to dataset-specific graphon structure on several datasets; that empirical cut-distance and cut-norm variance decreases with $n$ consistent with our bounds; and that attention graphons transfer to larger graph sizes with error decreasing in $n$.
Problem

Research questions and friction points this paper is trying to address.

Graph Transformers
Attention Graphons
Dense Graph Limit Theory
Cut-distance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Attention Graphons
Graph Transformers
Dense Graph Limit Theory
Cut-distance
Nonparametric Estimation
🔎 Similar Papers
2024-07-13arXiv.orgCitations: 36