Score
Designs and implements vector-quantized latent compression systems that map continuous representations to discrete codebook symbols and reconstruct them with a decoder (VQ‑VAE and related learned-codebook models). Builds and evaluates learned priors over those discrete symbols (including autoregressive priors) and integrates entropy coding based on the learned symbol distributions to produce compact bitstreams that exploit non‑uniform (e.g., power‑law) code usage.
This work addresses the challenge of optimizing rate-distortion trade-offs in video compression below 0.1 bits per pixel (bpp), where existing methods struggle due to the absence of differentiable rate signals. The authors propose MS-VQ-VAE, a framework leveraging discrete latent variables and vector quantization combined with an autoregressive prior to model codebook usage distributions, enabling ultra-low-bitrate compression without explicit rate penalties. By revealing that codebook utilization follows a power-law distribution, they employ entropy coding to push empirical bitrates below theoretical limits. To stabilize training with small codebooks, they introduce exponential moving average (EMA) codebook updates and a dead-code revival mechanism, mitigating gradient collapse. On UCF101, the method achieves 0.043–0.064 bpp—5× to 7.6× more efficient than H.265—and consistently surpasses H.265 (CRF=36) in perceptual quality measured by LPIPS, with gains up to 0.072.
Existing vector quantization (VQ) methods suffer from two key challenges: training instability—caused by gradient mismatch due to the straight-through estimator (STE)—and codebook collapse, stemming from poor codevector utilization. Both issues arise fundamentally from misalignment between the input feature distribution and the codebook distribution. This paper proposes a Wasserstein-distance-based distribution alignment optimization framework—the first to incorporate the Wasserstein distance directly into the VQ objective—to jointly mitigate gradient mismatch and underutilization from a distribution-matching perspective. We provide theoretical guarantees on convergence and quantization error bounds, and design a differentiable distribution alignment mechanism alongside an improved STE. Experiments demonstrate near-perfect codebook utilization (~100%), significantly reduced quantization error, and consistent improvements in reconstruction quality and training stability across multiple autoregressive modeling tasks.
In vector quantized variational autoencoders (VQ-VAEs), the vector quantization (VQ) operation is inherently non-differentiable, and conventional straight-through estimation (STE) introduces gradient distortion and information loss. To address this, we propose RotNorm-VQ: the first method to integrate a rotation-plus-normalization linear transformation into the VQ layer, establishing a smooth, differentiable mapping from encoder outputs to codebook vectors—enabling end-to-end gradient propagation through quantization. By parameterizing rotation via an orthogonal matrix, RotNorm-VQ preserves angular structure; combined with magnitude normalization, it retains amplitude information while avoiding the hard thresholding bias of STE. Evaluated across 11 mainstream VQ-VAE training paradigms, RotNorm-VQ consistently improves reconstruction quality (PSNR/SSIM ↑), codebook utilization (+23.6%), and quantization accuracy (quantization error ↓18.4%). This work establishes a novel differentiable vector quantization paradigm grounded in geometrically principled, continuous relaxation.
Vector quantization (VQ) in unsupervised learning often suffers from representation collapse, leading to low codebook utilization and degenerate latent spaces, thereby limiting model scalability. This paper proposes SimVQ, a novel VQ variant that reparameterizes the entire codebook via a learnable linear layer, shifting the optimization objective from selecting a single nearest-codebook vector to projecting onto the linear subspace spanned by the codebook. Designed through rigorous theoretical analysis, SimVQ integrates seamlessly into standard VQ frameworks without requiring auxiliary regularization or dimensionality reduction. Evaluated on multimodal image and audio tasks, SimVQ introduces only a lightweight linear transformation yet achieves substantial improvements in codebook utilization and downstream performance while effectively mitigating collapse. The implementation is publicly available.
Existing vector quantization (VQ) generative models rely on fixed codebooks, resulting in inflexible bitrates, the need for repeated retraining, and a fundamental trade-off between compression efficiency and reconstruction fidelity. This work proposes a multi-rate codebook adaptation framework that, for the first time, enables a single pre-trained VQ model to generate discrete representations at arbitrary bitrates without retraining. Our approach comprises two key innovations: (1) a data-driven mechanism for generating multi-rate codebooks, and (2) a lightweight adaptation method for pre-trained VQ models, leveraging hierarchical clustering and codebook embedding interpolation. Experiments demonstrate consistent and significant improvements over fixed-codebook baselines across diverse bitrates. The framework supports continuous, fine-grained rate-distortion control, substantially enhancing the generalizability, deployment flexibility, and inference efficiency of VQ models in practical applications.
This work addresses the instability and codebook collapse in vector quantization caused by distributional mismatch between features and the codebook by introducing distribution matching as a central principle and proposing a unified theoretical framework. The approach explicitly aligns the two distributions using either the Wasserstein distance—admitting a closed-form solution under Gaussian approximation—or a non-parametric maximum mean discrepancy (MMD). This alignment significantly enhances codebook utilization and stabilizes training. Experimental results demonstrate that the proposed method substantially outperforms existing approaches on visual tokenization benchmarks, exhibiting strong effectiveness, robustness, and efficient codebook usage.
本文提出了一种树结构矢量量化框架Tree-VQ,用于解决图像压缩中渐进式编码问题,通过组织离散码字为层次二叉树实现有效且渐进的图像压缩。
This work addresses the challenges of codebook collapse and instability in end-to-end training when applying vector quantization (VQ) to neural network weight compression. To mitigate these issues, the authors propose a cosine similarity–based codeword assignment strategy, combined with top-1 sampling and a straight-through estimator (STE) to enable stable and efficient VQ quantization. Furthermore, they integrate differentiable neural architecture search (NAS) to automatically determine an optimal per-layer configuration of mixed vector and linear quantization. By replacing conventional Euclidean distance with cosine similarity, the method avoids reconstruction bias caused by weighted averaging, thereby enhancing training stability without compromising compression efficiency. Although it does not uniformly outperform existing approaches across all quantization levels, the study offers critical insights into the underlying mechanisms and design trade-offs in VQ-based compression.
This work identifies and addresses the "entropy cliff" phenomenon in conventional discrete visual autoregressive models, where a fixed codebook causes a sharp drop in conditional entropy toward the end of the sequence, reducing generation to mere memorization and limiting reconstruction fidelity. To overcome this fundamental limitation, the authors propose Variable Codebook Quantization (VCQ), which monotonically increases codebook capacity along the sequence—from a minimum size \(K_{\text{min}} = 2\) to a maximum \(K_{\text{max}}\)—within a standard autoregressive Transformer architecture. Notably, VCQ requires no modifications to the loss function, model parameters, or training protocol, yet induces a coarse-to-fine semantic hierarchy. On ImageNet at 256×256 resolution, the base model reduces gFID from 27.98 to 14.80, with an extended variant achieving 1.71; furthermore, a linear probe using only the first 10 tokens attains 43.8% top-1 accuracy, surpassing the information-theoretic bottleneck imposed by fixed codebooks.
Existing vector quantization methods are constrained by static codebooks, limiting their ability to adapt to the heterogeneous geometric structures of data, while dynamic quantizers often suffer from inefficient serial decoding. This work proposes RQ-MoE, a novel framework that introduces the mixture-of-experts (MoE) mechanism into residual quantization for the first time. By employing a two-layer MoE module within a dual-stream architecture, RQ-MoE enables input-adaptive dynamic codebook construction and decouples token generation from quantization to facilitate parallel decoding. The framework unifies standard residual quantization and QINCo as special cases and provides design guidelines for expert dimensionality. Experiments demonstrate that RQ-MoE achieves state-of-the-art or comparable performance in reconstruction and retrieval tasks while delivering 6–14× faster decoding speeds than existing approaches.