🤖 AI Summary
This study addresses the spectral underfitting and overfitting issues in student models during GNN-to-MLP distillation, which arise from neglecting graph geometric structures. To overcome this limitation, we propose G²MLP, a novel framework that pioneers the incorporation of Ollivier-Ricci curvature into the distillation process to precisely localize error-prone regions. Based on this geometric insight, the method dynamically allocates supervision signals at both prediction and representation levels, guiding standard MLPs toward graph-free inference through energy-weighted teacher-student alignment. Our approach significantly narrows the rank gap between teacher and student models while improving node classification accuracy. Furthermore, it demonstrates strong generalization capabilities by successfully extending to Graph Transformers and link prediction tasks.
📝 Abstract
GNN-to-MLP distillation aims to retain the predictive accuracy of a message-passing teacher while deploying a graph-free MLP at inference. Existing methods mainly transfer node-wise predictions or use confidence-based reweighting, but they do not specify where the student should preserve the teacher's graph-induced geometry. We show that this omission leads to two spectral failure modes in the student's representation space. On sparse graphs, the student suffers from spectral underfit, missing high-energy teacher directions concentrated near boundary regions. On dense graphs, it suffers from spectral overfit, retaining spurious directions that the teacher has collapsed through aggregation. Motivated by an energy-weighted teacher-student alignment objective, we propose Graph Geometry-aware MLP (G^2MLP), a training-time distillation framework guided by Ollivier-Ricci curvature. Curvature identifies where the two spectral errors concentrate and is used to allocate supervision between prediction-level and representation-level alignment. The deployed model remains a standard MLP and requires no graph access at inference. Across node-classification benchmarks, G^2MLP consistently improves over graph-free distillation baselines, reduces the teacher-student rank gap in both regimes, and transfers without architectural changes to Graph Transformer teachers and link prediction.