๐ค AI Summary
This work addresses the challenges of codebook collapse and instability in end-to-end training when applying vector quantization (VQ) to neural network weight compression. To mitigate these issues, the authors propose a cosine similarityโbased codeword assignment strategy, combined with top-1 sampling and a straight-through estimator (STE) to enable stable and efficient VQ quantization. Furthermore, they integrate differentiable neural architecture search (NAS) to automatically determine an optimal per-layer configuration of mixed vector and linear quantization. By replacing conventional Euclidean distance with cosine similarity, the method avoids reconstruction bias caused by weighted averaging, thereby enhancing training stability without compromising compression efficiency. Although it does not uniformly outperform existing approaches across all quantization levels, the study offers critical insights into the underlying mechanisms and design trade-offs in VQ-based compression.
๐ Abstract
In this work, we developed and tested 3 techniques for vector quantization (VQ) based model weight compression. To mitigate codebook collapse and enable end-to-end training, we adopted cosine similarity-based assignment. Building on ideas from attention-based formulations in Differentiable K-Means (DKM), we further improved this approach by using cosine similarity for assignment combined with top-1 sampling and a straight-through estimator, thereby eliminating the need for weighted-average reconstruction. Finally, we investigated the use of differentiable neural architecture search (NAS) to adaptively select layer-wise quantization configurations, further optimizing the compression process. Although our method does not consistently outperform existing approaches across all quantization levels, it provides useful insights into the design trade-offs and behaviors of VQ-based model compression methods.