🤖 AI Summary
Existing studies lack quantitative theoretical analysis of the differences in optimization and generalization performance between graph neural networks (GNNs) and multilayer perceptrons (MLPs).
Method: Leveraging feature learning theory, this paper provides the first rigorous analysis of two-layer graph convolutional networks (GCNs) trained via gradient descent, modeling ReLU^q activation functions and spectral properties of the expected degree matrix D, under a structure-guided signal-noise separation framework.
Contribution/Results: We theoretically prove that graph convolution explicitly exploits graph structure to significantly enhance signal learning while suppressing noise memorization, thereby expanding the benign overfitting regime by approximately √D^{q−2} compared to CNNs. This constitutes the first quantitative characterization of the fundamental generalization advantage of GNNs over MLPs. Extensive simulations corroborate the superior generalization and robustness of GCNs predicted by our theory.
📝 Abstract
Graph neural networks (GNNs) have pioneered advancements in graph representation learning, exhibiting superior feature learning and performance over multilayer perceptrons (MLPs) when handling graph inputs. However, understanding the feature learning aspect of GNNs is still in its initial stage. This study aims to bridge this gap by investigating the role of graph convolution within the context of feature learning theory in neural networks using gradient descent training. We provide a distinct characterization of signal learning and noise memorization in two-layer graph convolutional networks (GCNs), contrasting them with two-layer convolutional neural networks (CNNs). Our findings reveal that graph convolution significantly augments the benign overfitting regime over the counterpart CNNs, where signal learning surpasses noise memorization, by approximately factor $sqrt{D}^{q-2}$, with $D$ denoting a node's expected degree and $q$ being the power of the ReLU activation function where $q>2$. These findings highlight a substantial discrepancy between GNNs and MLPs in terms of feature learning and generalization capacity after gradient descent training, a conclusion further substantiated by our empirical simulations.