Efficient VQ-QAT and Mixed Vector/Linear quantized Neural Networks

๐Ÿ“… 2026-04-25
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenges of codebook collapse and instability in end-to-end training when applying vector quantization (VQ) to neural network weight compression. To mitigate these issues, the authors propose a cosine similarityโ€“based codeword assignment strategy, combined with top-1 sampling and a straight-through estimator (STE) to enable stable and efficient VQ quantization. Furthermore, they integrate differentiable neural architecture search (NAS) to automatically determine an optimal per-layer configuration of mixed vector and linear quantization. By replacing conventional Euclidean distance with cosine similarity, the method avoids reconstruction bias caused by weighted averaging, thereby enhancing training stability without compromising compression efficiency. Although it does not uniformly outperform existing approaches across all quantization levels, the study offers critical insights into the underlying mechanisms and design trade-offs in VQ-based compression.

Technology Category

Machine Learning: Learning on the Edge & Model CompressionComputer Vision: Adversarial Attacks & RobustnessCognitive Modeling & Cognitive Systems: Neural Spike Coding

Application Category

Graph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsWeb Mining and Content Analysis: Robustness and generalizability of Web mining methods
๐Ÿ“ Abstract
In this work, we developed and tested 3 techniques for vector quantization (VQ) based model weight compression. To mitigate codebook collapse and enable end-to-end training, we adopted cosine similarity-based assignment. Building on ideas from attention-based formulations in Differentiable K-Means (DKM), we further improved this approach by using cosine similarity for assignment combined with top-1 sampling and a straight-through estimator, thereby eliminating the need for weighted-average reconstruction. Finally, we investigated the use of differentiable neural architecture search (NAS) to adaptively select layer-wise quantization configurations, further optimizing the compression process. Although our method does not consistently outperform existing approaches across all quantization levels, it provides useful insights into the design trade-offs and behaviors of VQ-based model compression methods.
Problem

Research questions and friction points this paper is trying to address.

vector quantization
model compression
codebook collapse
quantization
neural networks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vector Quantization
Cosine Similarity
Straight-Through Estimator
Differentiable NAS
Model Compression
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
T
Tianrun Gou
P
Puneet Gupta