Score
Design and implement methods that compress high-dimensional embedding vectors using product quantization (PQ): decompose each embedding into sub-vectors, learn shared codebooks for those subspaces, and encode sub-vectors as compact indices. Build the encoding/decoding pipeline so approximate embeddings can be reconstructed at inference, reducing storage and trainable embedding parameters.
To address the low reconstruction fidelity and inefficiency of latent representations in high-fidelity image generation, this paper proposes PQGAN—the first generative framework integrating product quantization (PQ) into latent variable encoding within the VQGAN architecture. PQGAN jointly leverages subspace decomposition and codebook optimization to enable efficient, high-fidelity quantization of high-dimensional latent spaces. We empirically identify an inverse relationship between embedding dimensionality and the relative performance of vector quantization versus product quantization, thereby establishing principled guidelines for hyperparameter selection. Experiments on ImageNet demonstrate that PQGAN achieves 37 dB PSNR—improving upon baseline VQGAN by 10 dB—while reducing FID, LPIPS, and CMMD by up to 96%. Moreover, PQGAN supports either doubling output resolution or accelerating generation, and seamlessly integrates with diffusion models.
Storing and retrieving high-dimensional embedding vectors incurs substantial memory overhead and computational cost. To address this, we propose Non-uniform Vector Quantization (NVQ), the first method that learns a customizable, lightweight nonlinear transformation for each index vector to enable personalized non-uniform quantization. Departing from conventional uniform quantization schemes, NVQ models local data distributions via compact, differentiable functions, achieving high-fidelity compression with minimal computational overhead. Extensive experiments on standard benchmarks demonstrate that NVQ consistently outperforms state-of-the-art methods—including PQ, OPQ, and AQ—at equivalent compression ratios, improving average Recall@10 by 3.2–7.8 percentage points while maintaining millisecond-scale query latency. Our core contributions are: (i) the first individualized non-uniform quantization framework tailored for approximate nearest neighbor (ANN) search; (ii) an efficient, learnable nonlinear transformation mechanism; and (iii) joint optimization of accuracy, efficiency, and compression ratio.
This work addresses the substantial storage and computational overhead of high-dimensional text embedding models by systematically investigating a joint compression strategy that combines dimensionality reduction and quantization. It is the first to demonstrate that the synergy between these two techniques can significantly outperform either approach applied in isolation. The effectiveness of the proposed method is validated across four MTEB task families and four widely used pretrained embedding models. Experimental results show that, in certain scenarios, the joint approach can reduce embedding size to as little as 0.1% of the original with negligible performance degradation, while the optimal compression strategy varies across tasks. This method thus offers a flexible and efficient solution for practical deployment, substantially lowering resource consumption without compromising embedding quality.
Vector quantization (VQ) in unsupervised learning often suffers from representation collapse, leading to low codebook utilization and degenerate latent spaces, thereby limiting model scalability. This paper proposes SimVQ, a novel VQ variant that reparameterizes the entire codebook via a learnable linear layer, shifting the optimization objective from selecting a single nearest-codebook vector to projecting onto the linear subspace spanned by the codebook. Designed through rigorous theoretical analysis, SimVQ integrates seamlessly into standard VQ frameworks without requiring auxiliary regularization or dimensionality reduction. Evaluated on multimodal image and audio tasks, SimVQ introduces only a lightweight linear transformation yet achieves substantial improvements in codebook utilization and downstream performance while effectively mitigating collapse. The implementation is publicly available.
To address the longstanding trade-off between model size and accuracy in large language model (LLM) quantization, this paper proposes GPTVQ, a high-dimensional vector quantization method. Methodologically, GPTVQ introduces four key innovations: (i) theoretical and empirical validation that increasing quantization dimensionality significantly improves the size–accuracy trade-off; (ii) Hessian-weighted output reconstruction loss for more accurate gradient-aware optimization; (iii) Hessian-guided column-wise alternating quantization to preserve layer-wise sensitivity; and (iv) data-aware EM initialization combined with integer-only quantization and SVD-based low-rank compression. Evaluated on Llama-2, Mistral, and other state-of-the-art LLMs, GPTVQ sets new post-training quantization SOTA: it quantizes 70B-parameter models on a single H100 GPU in just 3–11 hours, and achieves lower VQ decompression latency on mobile devices than standard 4-bit integer quantization—marking simultaneous advances in both accuracy retention and deployment efficiency.
This work addresses the challenges of codebook collapse and instability in end-to-end training when applying vector quantization (VQ) to neural network weight compression. To mitigate these issues, the authors propose a cosine similarity–based codeword assignment strategy, combined with top-1 sampling and a straight-through estimator (STE) to enable stable and efficient VQ quantization. Furthermore, they integrate differentiable neural architecture search (NAS) to automatically determine an optimal per-layer configuration of mixed vector and linear quantization. By replacing conventional Euclidean distance with cosine similarity, the method avoids reconstruction bias caused by weighted averaging, thereby enhancing training stability without compromising compression efficiency. Although it does not uniformly outperform existing approaches across all quantization levels, the study offers critical insights into the underlying mechanisms and design trade-offs in VQ-based compression.
Existing vector quantization methods are constrained by static codebooks, limiting their ability to adapt to the heterogeneous geometric structures of data, while dynamic quantizers often suffer from inefficient serial decoding. This work proposes RQ-MoE, a novel framework that introduces the mixture-of-experts (MoE) mechanism into residual quantization for the first time. By employing a two-layer MoE module within a dual-stream architecture, RQ-MoE enables input-adaptive dynamic codebook construction and decouples token generation from quantization to facilitate parallel decoding. The framework unifies standard residual quantization and QINCo as special cases and provides design guidelines for expert dimensionality. Experiments demonstrate that RQ-MoE achieves state-of-the-art or comparable performance in reconstruction and retrieval tasks while delivering 6–14× faster decoding speeds than existing approaches.
This work addresses a key limitation of conventional quantization methods, which optimize for minimal mean squared error yet often fail to preserve the inner products between vectors and arbitrary inputs, thereby degrading downstream task performance. The authors propose a novel quantization objective centered explicitly on inner product preservation, uncovering its intrinsic connection to Adaptive Stochastic Quantization (ASQ). Building on this insight, they design an unbiased, adaptive, and computationally efficient quantization algorithm that provably approximates inner product structures under both worst-case and average-case scenarios. The method offers strong theoretical guarantees while demonstrating practical efficacy across diverse data distributions. Empirically, their ASQ implementation achieves 2–10× speedups over current state-of-the-art approaches without sacrificing quantization accuracy.
Existing post-training quantization methods for large language models struggle to achieve efficient sub-1-bit compression, often hindered by high data or computational demands and additional storage overhead. This work proposes NanoQuant, the first post-training quantization framework capable of both binarization and sub-1-bit compression. By integrating low-rank binary matrix decomposition, ADMM-based initialization, and block-wise reconstruction, NanoQuant compresses Llama2-70B by 25.8× (averaging <1 bit per parameter) in under 13 hours on a single H100 GPU, enabling its deployment on an 8GB consumer-grade GPU. This dramatically lowers the hardware barrier for inference while establishing a new Pareto frontier between model accuracy and compression ratio.
This work addresses the instability and codebook collapse in vector quantization caused by distributional mismatch between features and the codebook by introducing distribution matching as a central principle and proposing a unified theoretical framework. The approach explicitly aligns the two distributions using either the Wasserstein distance—admitting a closed-form solution under Gaussian approximation—or a non-parametric maximum mean discrepancy (MMD). This alignment significantly enhances codebook utilization and stabilizes training. Experimental results demonstrate that the proposed method substantially outperforms existing approaches on visual tokenization benchmarks, exhibiting strong effectiveness, robustness, and efficient codebook usage.