Stable Transformers for Graph Generation

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the representation collapse and vanishing gradients in deep graph Transformers caused by repeated self-attention, which constrain generative performance. By analyzing the spectral dynamics of denoisers from a dynamical systems perspective, this work reveals that standard architectures inherently become dissipative as depth increases. Accordingly, it constructs a permutation-equivariant model featuring non-dissipative transport and introduces a tunable damping mechanism to isolate and evaluate dissipative effects. Experimental results demonstrate that non-dissipative dynamics effectively preserve representational diversity and maintain stable gradient flow, significantly outperforming highly dissipative counterparts. These findings establish non-dissipative dynamics as a critical design principle for deep graph models.
πŸ“ Abstract
Graph generative models increasingly rely on Graph Transformers (GT) to capture complex dependencies among nodes and edges. While deeper architectures should provide greater expressive capacity and a broader receptive field, their effectiveness can decline with depth: repeated self-attention progressively contracts node representations, impeding information flow and gradient propagation. We analyse this phenomenon from a dynamical systems perspective, focusing on how the denoiser's spectral dynamics affect graph generation. We show that standard GT denoisers become increasingly dissipative as depth grows, leading to vanishing gradients and representation collapse. To isolate the effect of these dynamics, we construct a permutation-equivariant GT with inherently stable, non-dissipative transport. We also introduce a damping mechanism that continuously interpolates between non-dissipative and increasingly contractive regimes, enabling a direct assessment of how dissipation influences generation. Experiments on synthetic and molecular graph generation benchmarks show that the gap between these regimes widens with depth: non-dissipative dynamics preserve representation diversity and gradient flow, sustaining strong generative performance, whereas greater contraction progressively impairs it. These findings identify the denoiser's dynamical regime as a key design factor for deep graph generative models.
Problem

Research questions and friction points this paper is trying to address.

Graph Generation
Graph Transformers
Representation Collapse
Vanishing Gradients
Dissipative Dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Graph Transformers
Dynamical Systems
Non-dissipative Transport
Permutation Equivariance
Damping Mechanism
πŸ”Ž Similar Papers
2024-07-13arXiv.orgCitations: 36