🤖 AI Summary
This study addresses the challenges of overfitting due to scarce supervision and representation drift caused by cross-session edges in graph few-shot class-incremental learning. To this end, we propose a replay-free framework that employs a frozen graph backbone combined with analytic continual learning to prevent catastrophic forgetting. Furthermore, we introduce a novel integration of Potts Markov random field inference to inject topological priors for optimizing sparse supervision, alongside a drift-aware knowledge distillation mechanism to mitigate cross-session representation shifts. Experimental results demonstrate that the proposed method significantly outperforms existing baselines across five benchmark datasets. Under the 5-shot setting, it achieves an average accuracy improvement exceeding 5%, reduces performance degradation by nearly 11%, and substantially enhances training efficiency.
📝 Abstract
Graph few-shot class-incremental learning (GFSCIL) requires a model to continually recognize emerging classes from only a few labeled nodes while preserving previously acquired knowledge. Beyond the catastrophic forgetting inherited from conventional graph continual learning, GFSCIL presents two distinctive challenges: extremely limited novel-class supervision causes severe overfitting, while cross-session edges---edges connecting newly arriving nodes with historical nodes---alter historical propagation neighborhoods and thereby induce representation drift. We propose MAGIC, a replay-free GFSCIL framework that combines a frozen graph representation backbone (e.g., an intrinsically parameter-free backbone such as SGC or a pretrained graph foundation model) with closed-form analytic continual learning. To alleviate novel-session overfitting, MAGIC learns a topological prior from the base graph that can characterize both homophilous and heterophilous relations, and injects this prior through Potts Markov random field inference to refine supervision for novel classes. To mitigate representation drift, MAGIC transfers previous predictions from the old representations of affected historical nodes to their updated representations through drift-aware analytic distillation. Experiments across five datasets and eight baselines demonstrate the effectiveness of MAGIC. Under the 5-shot setting, MAGIC improves Mean Accuracy and Final Accuracy by 5.48 percentage points and 9.33 percentage points on average, and reduces Performance Drop by 10.78 percentage points on average compared with the best baselines. MAGIC also shows clear advantages under the 1- and 3-shot settings, with larger gains as the number of supports increases. Moreover, MAGIC requires substantially less training time.