Transferable Graph Metanetworks

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of weight-space networks generalizing across varying widths by proposing Transferable Graphlet Networks. Methodologically, it integrates graph neural networks with Maximal Update Parametrization (μP), introduces functionally equivalent invariant constraints, and establishes a theoretical analysis framework in the infinite-width limit to ensure width invariance and continuity. The core contribution lies in enabling a cross-size transfer paradigm that permits training on small networks while predicting for larger ones. Experimental results demonstrate that under μP, the proposed method achieves robust generalization across up to 42-fold width variations, significantly improving cross-scale performance prediction accuracy. This work provides both theoretical foundations and methodological support for neural network architecture evaluation.
📝 Abstract
A weight space network (or metanetwork) takes the weights of another neural network as input and predicts properties of it. Most prior work trains such models on input networks of one or a few fixed sizes and evaluates them in-distribution. The few attempts at out-of-distribution size generalization remain limited in scope and have achieved only modest success. Consequently, the potential efficiency gains of training on small networks and evaluating on much larger ones remain largely unrealized. We propose Transferable Graph Metanetworks, which extend the graph metanetwork paradigm with a set of modifications that make performance transferable across input networks of different widths. The modifications follow two principles: invariance to the ways in which networks of different widths represent the same function, and continuity, such that weights representing similar functions receive similar predictions. We further study whether size generalization is possible for input networks trained independently from random initialization. Empirically, our modifications significantly improve size generalization on every task we consider. Performance is strongest on input networks trained under the maximal-update parameterization ($μ$P), where it remains robust up to $42\times$ the training width. Theoretically, we explain these observations with infinite-width limit theory: we prove size-generalization guarantees for our model on $μ$P-trained inputs, and explain why it can fail under other parameterizations.
Problem

Research questions and friction points this paper is trying to address.

weight space networks
metanetworks
size generalization
out-of-distribution
transferability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Graph Metanetworks
Size Generalization
Weight Space Networks
Maximal-Update Parameterization
Infinite-Width Limit Theory
🔎 Similar Papers
2024-02-16Nature CommunicationsCitations: 2