🤖 AI Summary
This work addresses the high storage and transmission costs of neural network parameters by proposing an extreme compression method based on random seeds and quantized latent variables. The model weights are represented as mappings generated jointly by a fixed random basis, trainable quantized latent vectors, and seed-based initialization. This approach eliminates the need to store full weight matrices or projection matrices, requiring only a compact set of latent vectors for model reconstruction, while preserving accuracy through quantization-aware fine-tuning. Leveraging a block-wise scalable basis design, the method efficiently supports deployment of large-scale models. Experiments demonstrate that, under comparable accuracy to low-bitwidth quantized models, the compressed model size is drastically reduced—determined solely by the dimensionality and bitwidth of the latent variables rather than the original parameter count.
📝 Abstract
The cost of storing and transmitting a trained neural network scales with its parameter count, a bottleneck for over-the-air updates, on-device libraries, and other bandwidth-bound deployments. We study an extreme form of model compression in which the deployable artifact is not the weights but a short recipe for regenerating them. Building on Mapping Networks, which express a network's weights as a nonlinear function of a compact trainable latent and a fixed random basis, we observe that only the latent need be stored, because the basis and initialization center are reproducible from an integer seed. A model becomes a seed together with a quantized latent, whose size is set by the latent dimension and bit width rather than the parameter count. We formalize this artifact and introduce a seeded block-wise basis that scales to networks whose projection cannot be held in memory. In our experiments, a mapped model is as accurate as the same network quantized aggressively to a few bits per weight, while taking far fewer bytes to store. Reaching the most aggressive bit widths depends on fine-tuning the latent with quantization in the loop. The results do not depend on the particular random basis, and a structured basis lets the weights be regenerated almost for free even for large networks.