Graph Neural Networks Provably Benefit from Structural Information: A Feature Learning Perspective

📅 2023-06-24
🏛️ arXiv.org
📈 Citations: 8
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing studies lack quantitative theoretical analysis of the differences in optimization and generalization performance between graph neural networks (GNNs) and multilayer perceptrons (MLPs). Method: Leveraging feature learning theory, this paper provides the first rigorous analysis of two-layer graph convolutional networks (GCNs) trained via gradient descent, modeling ReLU^q activation functions and spectral properties of the expected degree matrix D, under a structure-guided signal-noise separation framework. Contribution/Results: We theoretically prove that graph convolution explicitly exploits graph structure to significantly enhance signal learning while suppressing noise memorization, thereby expanding the benign overfitting regime by approximately √D^{q−2} compared to CNNs. This constitutes the first quantitative characterization of the fundamental generalization advantage of GNNs over MLPs. Extensive simulations corroborate the superior generalization and robustness of GCNs predicted by our theory.
📝 Abstract
Graph neural networks (GNNs) have pioneered advancements in graph representation learning, exhibiting superior feature learning and performance over multilayer perceptrons (MLPs) when handling graph inputs. However, understanding the feature learning aspect of GNNs is still in its initial stage. This study aims to bridge this gap by investigating the role of graph convolution within the context of feature learning theory in neural networks using gradient descent training. We provide a distinct characterization of signal learning and noise memorization in two-layer graph convolutional networks (GCNs), contrasting them with two-layer convolutional neural networks (CNNs). Our findings reveal that graph convolution significantly augments the benign overfitting regime over the counterpart CNNs, where signal learning surpasses noise memorization, by approximately factor $sqrt{D}^{q-2}$, with $D$ denoting a node's expected degree and $q$ being the power of the ReLU activation function where $q>2$. These findings highlight a substantial discrepancy between GNNs and MLPs in terms of feature learning and generalization capacity after gradient descent training, a conclusion further substantiated by our empirical simulations.
Problem

Research questions and friction points this paper is trying to address.

Quantify GNNs' optimization and generalization advantages over MLPs
Compare signal learning and noise memorization in GCNs vs MLPs
Analyze feature learning theory to explain GNNs' superior performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

GNNs prioritize signal learning over noise
Feature learning theory analyzes GNN-MLP differences
Quantifies GNN advantage as D^(q-2) times
🔎 Similar Papers
No similar papers found.
RIKEN Center for Advanced Intelligence Project | The University of Hong Kong | The University of New South Wales | The University of Tokyo
W
Wei Huang
RIKEN Center for Advanced Intelligence Project
Y
Yuan Cao
The University of Hong Kong
H
Hong Wang
X
Xin Cao
The University of New South Wales
Taiji Suzuki
Taiji Suzuki
The University of Tokyo
StatisticsMachine learning