🤖 AI Summary
This work systematically investigates the performance-efficiency trade-offs between dense and sparse attention mechanisms in graph transformers for modeling long-range node dependencies. We conduct controlled experiments across diverse graph structures—including social networks, molecular graphs, and citation networks—to quantitatively characterize their expressive power, computational cost, memory footprint, and scalability. Guided by empirical findings, we propose a “structure-aware sparsification” design principle: hybrid local-global sparse attention for high-clustering-coefficient graphs, and selective retention of global dense connections for small-diameter, homogeneous graphs. Experiments demonstrate that our strategy preserves over 95% of dense-attention accuracy while reducing inference latency by 3.2× and GPU memory consumption by 68%. The study reveals a strong coupling between graph topological properties and attention paradigm efficacy, establishing an evidence-based foundation for architecture search and hardware-aware optimization of efficient graph transformers.
📝 Abstract
Graphs have become a central representation in machine learning for capturing relational and structured data across various domains. Traditional graph neural networks often struggle to capture long-range dependencies between nodes due to their local structure. Graph transformers overcome this by using attention mechanisms that allow nodes to exchange information globally. However, there are two types of attention in graph transformers: dense and sparse. In this paper, we compare these two attention mechanisms, analyze their trade-offs, and highlight when to use each. We also outline current challenges and problems in designing attention for graph transformers.