Score
Design and implement vector quantization systems that partition input vectors into clusters and then learn or apply separate codebooks per cluster, including the cluster-assignment procedure, per-cluster codebook training, and encoding/decoding pipelines. Measure and optimize reconstruction distortion, compression rate (bits per vector/item), and codebook sizing/assignment to minimize distortion or bitrate for clustered data.
This work addresses the poor rate-distortion efficiency of high-fidelity 2D Gaussian image representations, which stems from their massive number of floating-point parameters. To tackle this issue, the authors propose a clustering-guided vector quantization method (CGVQ) that introduces, for the first time, a clustering-based grouping strategy to partition Gaussian parameters into homogeneous groups prior to quantization. This approach significantly reduces bitrate while preserving high reconstruction accuracy. Experimental results demonstrate that CGVQ achieves up to a 20% reduction in bitrate compared to existing baselines, while maintaining comparable visual quality, thereby substantially improving the compression efficiency of 2D Gaussian representations.
Existing vector quantization (VQ) methods suffer from two key challenges: training instability—caused by gradient mismatch due to the straight-through estimator (STE)—and codebook collapse, stemming from poor codevector utilization. Both issues arise fundamentally from misalignment between the input feature distribution and the codebook distribution. This paper proposes a Wasserstein-distance-based distribution alignment optimization framework—the first to incorporate the Wasserstein distance directly into the VQ objective—to jointly mitigate gradient mismatch and underutilization from a distribution-matching perspective. We provide theoretical guarantees on convergence and quantization error bounds, and design a differentiable distribution alignment mechanism alongside an improved STE. Experiments demonstrate near-perfect codebook utilization (~100%), significantly reduced quantization error, and consistent improvements in reconstruction quality and training stability across multiple autoregressive modeling tasks.
Adaptive vector quantization (AVQ) is essential for compressing gradients, weights, activations, and datasets in machine learning, yet existing methods suffer from prohibitive time and memory complexity. Method: We propose the first algorithms for AVQ—both strictly optimal and highly efficient near-optimal—overcoming these bottlenecks. Our optimal algorithm employs a progressive dynamic programming framework with greedy pruning and error-bounded divide-and-conquer. For large-scale inputs, we introduce a super-fast near-optimal variant leveraging structural approximations and rigorous error analysis. Contribution/Results: The optimal algorithm guarantees theoretical precision, while the near-optimal variant achieves controllable distortion with 10–100× speedup and significantly reduced memory footprint. Both support seamless end-to-end integration into modern ML systems, enabling practical AVQ deployment across training and inference pipelines.
To address the high memory consumption, low computational efficiency, and lack of theoretical convergence guarantees in traditional K-means–based clustering for high-dimensional data, this paper proposes a robust clustering framework based on Stochastic Quantization (SQ). We are the first to systematically integrate SQ—a method with strong theoretical convergence guarantees—into both unsupervised and semi-supervised clustering. The framework incorporates a Triplet Network to enable dimensionality-adaptive embedding and low-dimensional latent space modeling, while leveraging mini-batch optimization for scalability. Evaluated on partially labeled image classification tasks, our approach achieves significantly faster convergence and reduced memory footprint compared to K-means++ and mini-batch K-means, while attaining superior clustering performance. This demonstrates a unified alignment between theoretical convergence properties and practical efficacy.
Existing vector quantization (VQ) generative models rely on fixed codebooks, resulting in inflexible bitrates, the need for repeated retraining, and a fundamental trade-off between compression efficiency and reconstruction fidelity. This work proposes a multi-rate codebook adaptation framework that, for the first time, enables a single pre-trained VQ model to generate discrete representations at arbitrary bitrates without retraining. Our approach comprises two key innovations: (1) a data-driven mechanism for generating multi-rate codebooks, and (2) a lightweight adaptation method for pre-trained VQ models, leveraging hierarchical clustering and codebook embedding interpolation. Experiments demonstrate consistent and significant improvements over fixed-codebook baselines across diverse bitrates. The framework supports continuous, fine-grained rate-distortion control, substantially enhancing the generalizability, deployment flexibility, and inference efficiency of VQ models in practical applications.
This work addresses the instability and codebook collapse in vector quantization caused by distributional mismatch between features and the codebook by introducing distribution matching as a central principle and proposing a unified theoretical framework. The approach explicitly aligns the two distributions using either the Wasserstein distance—admitting a closed-form solution under Gaussian approximation—or a non-parametric maximum mean discrepancy (MMD). This alignment significantly enhances codebook utilization and stabilizes training. Experimental results demonstrate that the proposed method substantially outperforms existing approaches on visual tokenization benchmarks, exhibiting strong effectiveness, robustness, and efficient codebook usage.
This study addresses the embedding dimension redundancy and unclear learning mechanisms when Transformers perform K-means clustering. We propose compressing the embedding dimension to d+log k, constructing a minimal Transformer that efficiently executes Lloyd's algorithm. By integrating stochastic gradient optimization with probing techniques, we theoretically characterize the convergence properties and in-distribution generalization conditions of the learned algorithm. Our experiments validate the effectiveness of this streamlined model on clustering tasks and identify the critical factors governing its success and failure. Ultimately, this work provides both theoretical and empirical foundations for understanding the algorithmic learning mechanisms underlying Transformers.
This work addresses the lack of a unified theoretical framework in existing vector quantization methods, which often struggle to balance geometric fidelity and compression efficiency. We propose Block-Sphere Quantization (BlockQuant), a novel rotation-based block quantization paradigm grounded in spherical geometry. By applying spherical quantization to blocks of randomly rotated embedding vectors, BlockQuant more faithfully preserves the original geometric structure. Our theoretical analysis provides the first unified comparison of mainstream rotation-based quantizers, establishing BlockQuant’s superiority in terms of expected distortion and mean squared error. Empirical results demonstrate that BlockQuant consistently achieves significantly lower reconstruction error and higher inner product fidelity than state-of-the-art methods across real-world embedding datasets and long-context large language model inference tasks.
Existing vector quantization methods are constrained by static codebooks, limiting their ability to adapt to the heterogeneous geometric structures of data, while dynamic quantizers often suffer from inefficient serial decoding. This work proposes RQ-MoE, a novel framework that introduces the mixture-of-experts (MoE) mechanism into residual quantization for the first time. By employing a two-layer MoE module within a dual-stream architecture, RQ-MoE enables input-adaptive dynamic codebook construction and decouples token generation from quantization to facilitate parallel decoding. The framework unifies standard residual quantization and QINCo as special cases and provides design guidelines for expert dimensionality. Experiments demonstrate that RQ-MoE achieves state-of-the-art or comparable performance in reconstruction and retrieval tasks while delivering 6–14× faster decoding speeds than existing approaches.
This work addresses the substantial storage and computational overhead of high-dimensional text embedding models by systematically investigating a joint compression strategy that combines dimensionality reduction and quantization. It is the first to demonstrate that the synergy between these two techniques can significantly outperform either approach applied in isolation. The effectiveness of the proposed method is validated across four MTEB task families and four widely used pretrained embedding models. Experimental results show that, in certain scenarios, the joint approach can reduce embedding size to as little as 0.1% of the original with negligible performance degradation, while the optimal compression strategy varies across tasks. This method thus offers a flexible and efficient solution for practical deployment, substantially lowering resource consumption without compromising embedding quality.