🤖 AI Summary
Kolmogorov–Arnold Networks (KANs) exhibit strong expressivity by replacing edge weights with basis coefficient vectors; however, their parameters and memory overhead scale multiplicatively, severely hindering deployment. To address this, we propose MetaCluster—a meta-learning-driven compression framework tailored for KANs. It employs a lightweight meta-learner to induce coefficient vectors to concentrate on a low-dimensional manifold, thereby enabling natural compatibility with K-means clustering. After training, only a compact codebook and sparse indices are retained, while the meta-network is discarded, achieving amortized storage. Our method attains up to 80× parameter compression across multiple datasets and KAN variants—without accuracy loss—significantly enhancing deployability. The core innovation lies in embedding meta-learning into parameter structural modeling, enabling, for the first time, efficient clusterability and lossless deep compression of KAN coefficients.
📝 Abstract
Kolmogorov-Arnold Networks (KANs) replace scalar weights with per-edge vectors of basis coefficients, thereby boosting expressivity and accuracy but at the same time resulting in a multiplicative increase in parameters and memory. We propose MetaCluster, a framework that makes KANs highly compressible without sacrificing accuracy. Specifically, a lightweight meta-learner, trained jointly with the KAN, is used to map low-dimensional embedding to coefficient vectors, shaping them to lie on a low-dimensional manifold that is amenable to clustering. We then run K-means in coefficient space and replace per-edge vectors with shared centroids. Afterwards, the meta-learner can be discarded, and a brief fine-tuning of the centroid codebook recovers any residual accuracy loss. The resulting model stores only a small codebook and per-edge indices, exploiting the vector nature of KAN parameters to amortize storage across multiple coefficients. On MNIST, CIFAR-10, and CIFAR-100, across standard KANs and ConvKANs using multiple basis functions, MetaCluster achieves a reduction of up to 80$ imes$ in parameter storage, with no loss in accuracy. Code will be released upon publication.