CrossGMN: Graph Metanetworks for Cross-Architecture Weight-Space Transformations

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the failure of equivariance in cross-architecture weight space transformations caused by mismatched permutation symmetries between source and target networks. To overcome this, it proposes a dual-input formulation and CrossGMN, a graph meta-network that jointly processes both networks through a symmetry-preserving cross-network message passing mechanism. Furthermore, a general theoretical proof for continuous cross-architecture operators is established, enabling precise weight generation that is invariant to the source and equivariant to the target. This work constructs a symmetry-preserving weight space meta-learning framework. Empirically, it accelerates knowledge distillation by 8.89× in model compression tasks while supporting zero-retraining cross-dataset transfer and unified compression across heterogeneous architectures.
📝 Abstract
Weight-space networks operate directly on parameters of other neural networks, enabling tasks such as predicting model properties, editing trained models, and generating weights. Weight-space symmetries such as neuron permutations make equivariance a key design principle. However, existing equivariant weight-space architectures have primarily been studied for transformations that preserve the network architecture. In contrast, many practical transformations, including model compression and upscaling, map a trained source network into a target network with a different architecture. In this setting, the source and target permutation symmetries act on different parameter spaces, making equivariance less straightforward to formulate. Our key idea for addressing this mismatch is to reformulate cross-architecture operators with two inputs: a trained source network and an initialization of the target network. This lets us define equivariant cross-architecture operators that refine the initialization of the target network using information from the source network, while being invariant to source-network permutations and equivariant to target-network permutations. Based on this formulation, we introduce CrossGMN, a graph metanetwork that jointly processes both networks through symmetry-preserving cross-network message passing. We prove CrossGMN is universal for continuous cross-architecture operators on compact sets under a general-position assumption. We evaluate CrossGMN for model compression, predicting a smaller network's parameters to accelerate subsequent knowledge distillation. Across 2-D and 3-D INRs and image classification with MLPs, CNNs, and Vision Transformers, CrossGMN speeds up distillation by up to 8.89x, transfers across datasets without retraining (3.78x), and a single model can accelerate compression from heterogeneous source architectures into a common target architecture.
Problem

Research questions and friction points this paper is trying to address.

weight-space networks
cross-architecture transformations
equivariance
permutation symmetry
model compression
Innovation

Methods, ideas, or system contributions that make the work stand out.

Graph Metanetworks
Cross-Architecture
Weight-Space Equivariance
Model Compression
Knowledge Distillation