Score
Designs and trains discrete codebooks and associated encoders that quantize continuous representations into compact indices (including spatial-index codebooks), optimizing quantization objectives under bit‑budget and compactness constraints. Builds evaluation and selection procedures that enforce response‑awareness and receiver separability—measuring reconstruction and downstream task fidelity, pruning or reformatting entries that produce ambiguous indices, and balancing task performance with compression.
This work addresses the challenge of maintaining perceptual quality consistency in learned video codecs when deployed across varying spatial resolutions, a scenario that typically necessitates retraining or rate-distortion parameter tuning. Building upon the MS-VQ-VAE framework, the study systematically investigates the impact of codebook capacity and spatial resolution on perceptual quality using the UCF101 dataset. The findings reveal that codebook capacity exerts an influence approximately ten times greater than that of resolution, with higher resolutions yielding superior entropy efficiency. These insights offer a novel perspective for designing discrete tokenizers in multi-resolution video compression and generative models. Experimental results demonstrate that the proposed method achieves LPIPS scores surpassing H.264 by 25–52% at 128×128 resolution and outperforming H.265 by 21–37% at 256×256, all while operating at comparable or lower bitrates.
To address low quantization efficiency, poor robustness, and semantic distortion induced by channel noise in digital semantic communication, this paper proposes an information-theoretic, learnable codebook design framework. We first establish a theoretical equivalence between semantic synonym mapping and Voronoi partitioning, then formulate an end-to-end jointly optimized objective comprising a semantic-maximizing entropy-regularized quantization loss and a channel-aware semantic distortion loss. The method integrates mutual information maximization, Voronoi-based quantization modeling, and channel distortion characterization. Evaluated on image reconstruction, the proposed approach achieves a 24.1% PSNR gain and a 46.5% improvement in LPIPS perceptual similarity at 10 dB SNR, significantly mitigating semantic distortion. This work introduces a novel paradigm for efficient and reliable semantic-driven transmission.
This work addresses the challenge of balancing accuracy and efficiency in large language model inference under memory constraints, particularly for massive models with mixture-of-experts architectures, where conventional quantization methods rely on handcrafted bit-width assignments, calibration data, or Hessian information. The authors propose XFP, a dynamic quantization framework that introduces a novel channel-wise cosine similarity–based adaptive mechanism requiring neither calibration nor Hessian computation. XFP employs an H-Process to automatically search for optimal configurations satisfying both quality and memory constraints, decomposing weights into sparse fp16 outlier residuals and dense sub-byte indices, while co-designing codebooks, outlier handling, and fused decoding kernels. On Qwen3.5-122B, XFP achieves 138 tokens/s and 94.49% GSM8K accuracy—49% faster than Marlin INT4. On Qwen3.5-397B, it fits within 2×96GB GPUs at ~3.4 effective bits, delivering 100.9 tokens/s in long-context decoding, outperforming INT4 with expert pruning in both accuracy and efficiency.
This work addresses the lack of systematic investigation into format selection and performance trade-offs in existing low-bit quantization-aware training (QAT) methods, as well as their insufficient evaluation on generative tasks. To this end, we propose the first integration of k-means clustering into QAT for 1-bit weight quantization, optimizing generative performance under a fixed inference memory budget. Our approach transcends the limitations of conventional integer-based quantization schemes by leveraging learned cluster centroids to better preserve model fidelity at ultra-low bitwidths. Experimental results demonstrate that, under identical memory constraints, our method significantly outperforms state-of-the-art integer quantization approaches while maintaining compatibility with general-purpose hardware for efficient deployment.
Existing global shared codebook approaches neglect intra-face semantic correlations and token-level semantic disparities, leading to suboptimal reconstruction quality and face recognition performance at ultra-low bitrates (e.g., 0.05 bpp). To address this, we propose a switchable token-specific codebook quantization framework: first, codebooks are learned independently per semantic category; then, each visual token is dynamically assigned its most suitable dedicated codebook, enabling fine-grained, low-distortion quantization. Our method is the first to jointly couple token-level codebook selection with category-aware grouping—reducing individual codebook size while enhancing representational diversity and quantization fidelity. Experiments demonstrate that reconstructed face images achieve a mean recognition accuracy of 93.51% at 0.05 bpp, significantly outperforming global codebook baselines. This work establishes a novel paradigm for codebook-driven face compression models.
This work addresses the unclear impact of the execution order between pruning and quantization in joint compression on model performance. It presents the first systematic investigation into this ordering effect and proposes the “progressive intensity hypothesis,” which posits that weaker perturbations should precede stronger ones—a claim substantiated through theoretical perturbation analysis. Extensive experiments across large language and vision models, including complex scenarios such as multi-stage compression and mixed-precision quantization, validate the universality of this hypothesis. The results demonstrate that adhering to the progressive intensity ordering consistently yields significant performance improvements and exhibits strong generalization across diverse architectures and compression configurations.
This work addresses the challenges of codebook collapse and instability in end-to-end training when applying vector quantization (VQ) to neural network weight compression. To mitigate these issues, the authors propose a cosine similarity–based codeword assignment strategy, combined with top-1 sampling and a straight-through estimator (STE) to enable stable and efficient VQ quantization. Furthermore, they integrate differentiable neural architecture search (NAS) to automatically determine an optimal per-layer configuration of mixed vector and linear quantization. By replacing conventional Euclidean distance with cosine similarity, the method avoids reconstruction bias caused by weighted averaging, thereby enhancing training stability without compromising compression efficiency. Although it does not uniformly outperform existing approaches across all quantization levels, the study offers critical insights into the underlying mechanisms and design trade-offs in VQ-based compression.
This work addresses the inconsistent performance of binary quantization in embedding spaces, which excels in contrastive learning embeddings but degrades sharply in others, and resolves the lack of a unified theoretical foundation between the “random rotation” and “axis-aligned” quantization strategies. The study identifies the heterogeneity of coordinate-wise variances as the key factor governing quantization efficacy and establishes, for the first time, an analytical framework under a Gaussian structural assumption. This framework yields a closed-form solution for rank fidelity, quantitatively linking the information content of magnitude bits to variance heterogeneity, and unifies the conditions under which the two seemingly opposing strategies are optimal. Theoretical predictions are validated across 13 datasets and 6 embedding types, providing the first principled design guidelines for binary quantization systems.