🤖 AI Summary
This study addresses the limited cross-dataset generalization of existing neural graph edit distance (GED) models, a performance gap obscured by conventional evaluation paradigms. To tackle this issue, the authors first quantify the cross-set transfer gap and propose a multi-source joint pretraining strategy leveraging exact GED supervision. Furthermore, they establish a comprehensive evaluation framework for zero-shot and few-shot transfer scenarios. The findings reveal significant performance degradation caused by discrepancies between training and test distributions. Crucially, the results demonstrate that multi-source pretraining substantially enhances target-domain performance and adaptability to novel tasks. By exposing these previously overlooked generalization limitations and validating an effective mitigation strategy, this work introduces a new paradigm for rigorously evaluating the generalizability of neural GED models across diverse domains.
📝 Abstract
Neural approaches to Graph Edit Distance (GED) have achieved strong results under standard within-dataset evaluation, but much less is known about how well these models transfer across graph collections. We conduct a systematic study of this problem using exact GED supervision across diverse graph datasets and a broad set of representative learning-based methods. Our results reveal a pronounced gap between within-collection performance and cross-collection transfer. Models that perform well on their training collections often lose this advantage when evaluated on structurally different data. Training on multiple source collections substantially improves zero-shot transfer and provides a better starting point when limited supervision is available for a new target collection. Further analysis shows that transfer behavior varies with the source--target direction and the structural characteristics of the collections involved. These findings suggest that conventional within-collection evaluation provides only a partial view of the generalization behavior of neural GED models and motivate broader evaluation across heterogeneous graph collections.