residual vector quantization

Designs and implements algorithms and pipelines that compress continuous high‑dimensional feature vectors or streams into compact discrete codes by staging residual quantizers, coordinate and lattice‑based quantizers (e.g., Z, A2, D4, E8 lattices), rotated or randomized‑basis binary/int4 quantizers, and tokenization/codebook components that emit indices and reconstruction decoders. Analyzes and optimizes the resulting encodings — rate‑distortion tradeoffs, decorrelation or rotation strategies, multi‑stage residual encoding, compact index representations for storage and fast lookup, and unbiased/bounded attention estimation using binary or int4 arithmetic — to preserve local/temporal or semantic structure while minimizing bitrate.

residualvectorquantization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing vector quantization methods are constrained by static codebooks, limiting their ability to adapt to the heterogeneous geometric structures of data, while dynamic quantizers often suffer from inefficient serial decoding. This work proposes RQ-MoE, a novel framework that introduces the mixture-of-experts (MoE) mechanism into residual quantization for the first time. By employing a two-layer MoE module within a dual-stream architecture, RQ-MoE enables input-adaptive dynamic codebook construction and decouples token generation from quantization to facilitate parallel decoding. The framework unifies standard residual quantization and QINCo as special cases and provides design guidelines for expert dimensionality. Experiments demonstrate that RQ-MoE achieves state-of-the-art or comparable performance in reconstruction and retrieval tasks while delivering 6–14× faster decoding speeds than existing approaches.

codebook adaptationdecoding bottleneckheterogeneous data geometry

Efficient Feature Compression for Machines with Global Statistics Preservation

Dec 09, 2025
ME
Md Eimran Hossain Eimon
🏛️ Florida Atlantic University | InterDigital - AI Lab

To address the bandwidth bottleneck in AI model split inference caused by intermediate feature transmission, this paper proposes a lightweight, lossless feature compression method based on Z-score normalization. The core innovation lies in the first integration of Z-score standardization into a feature compression framework to explicitly preserve global statistical properties—namely, mean and variance—enabling an end-to-end differentiable encoding architecture. Unlike the conventional scaling scheme in the MPEG FCM standard draft, our approach yields a more compact and hardware-efficient implementation. Extensive evaluation across multiple vision tasks demonstrates an average bitrate reduction of 17.09%, with up to 65.69% savings in object tracking, while maintaining zero accuracy degradation in downstream tasks.

Compress intermediate feature data in split AI inferenceImprove compression efficiency while preserving global statisticsReduce bitrate without sacrificing end-task accuracy

Q2D2: A Geometry-Aware Audio Codec Leveraging Two-Dimensional Quantization

Dec 01, 2025
EN
Eliya Nachmani
🏛️ Ben-Gurion University

Current neural audio codecs employ quantization schemes—such as residual vector quantization (RVQ), vector quantization (VQ), and finite scalar quantization (FSQ)—that suffer from limited geometric modeling capacity in latent space, resulting in weak correlation capture among features, low codebook utilization, and high token rates. To address this, we propose Q2D2, a geometry-aware audio compression framework that, for the first time, jointly quantizes feature pairs onto structured 2D grids (hexagonal, rhombic, or rectangular), implicitly constructing efficient codebooks. This approach abandons the oversimplified manifold assumptions of scalar or vector quantization, enhancing geometric consistency and collaborative feature representation while preserving reconstruction fidelity. Experiments on speech reconstruction demonstrate that Q2D2 matches or surpasses state-of-the-art models in both objective and subjective quality metrics, achieves significantly higher codebook utilization, and validates—through ablation—the critical role of 2D grid-based quantization.

Captures feature correlations better than traditional quantization methodsEnhances codebook utilization while maintaining reconstruction qualityImproves audio compression efficiency with low token rates

本文提出了一种几何感知的双曲残差量化方法,解决了双曲空间中残差聚合和梯度估计的几何不一致性问题,适用于层次化数据表示。

Geometric InconsistenciesHyperbolic GeometryResidual Vector Quantization

A Preprocessing Framework for Video Machine Vision under Compression

Dec 17, 2025
FZ
Fei Zhao
🏛️ Peking University | Bytedance

Existing video compression methods are primarily optimized for human visual perception and thus often fail to preserve semantic information critical for machine vision tasks. To address this, this paper proposes a machine-vision-oriented neural preprocessing framework. It introduces a learnable preprocessor prior to standard video encoding and pioneers a differentiable virtual codec, enabling end-to-end joint optimization of preprocessing and conventional encoders (e.g., H.264/AVC) without modifying codec standards. A rate–distortion–task loss jointly optimizes bit rate, reconstruction fidelity, and downstream task performance—including object detection and action recognition. Experiments demonstrate that the framework reduces average bit rate by over 15% while maintaining or even improving task accuracy, significantly enhancing semantic fidelity and utility of compressed video for machine vision applications.

Enhances rate-accuracy performance over human-centric metricsOptimizes video compression for machine vision tasksSaves bitrate with a practical preprocessing framework

Latest Papers

What's happening recently
View more

Discrete audio representations in speech language models often degrade downstream task performance due to information loss. To address this, this work proposes a hybrid discrete-continuous modeling approach that jointly represents speech using temporally compressed discrete tokens and dimensionality-reduced continuous residuals. The method introduces a novel encoder-decoder architecture incorporating fusion-focused modulation and a hybrid Transformer design, enabling autoregressive inference in the discrete domain while simultaneously leveraging non-autoregressive prediction and continuous residual upsampling. This approach achieves the first effective integration of discrete and continuous representations, substantially reducing the number of autoregressive steps while preserving speaker characteristics and fine-grained acoustic details. Experimental results demonstrate clear performance gains over purely discrete baselines.

discrete audio representationsinformation lossLarge Language Models

This study addresses the limitation that quantization and generative residuals impose on reconstruction quality in autoregressive image coding by proposing the ResARC framework. This method introduces a novel explicit dual residual compensation mechanism, leveraging diffusion models to compensate for quantization residuals while compressing and transmitting generative residuals. By integrating diffusion transformers with learned codecs, ResARC achieves high-fidelity context-based reconstruction without requiring additional side information. Experimental results demonstrate that ResARC significantly enhances distributional fidelity at ultra-low bitrates while maintaining highly competitive perceptual similarity.

autoregressive codinggeneration residualgenerative compression

This study addresses the optimization of high-precision coordinate retention strategies in rotation-based quantization. We propose a method that jointly optimizes the number and positions of retained coordinates before and after rotation to minimize quantization error under a fixed bit budget. Theoretically, we prove that retaining the top-k coordinates prior to rotation minimizes the upper bound of the quantization error, thereby reducing a complex combinatorial search to the optimization of a scalar k, for which a fast parallel selection algorithm is designed. Combined with random rotation preprocessing and offline codebook optimization, our approach achieves efficient compression. Experiments demonstrate that the proposed method significantly improves the trade-off between reconstruction accuracy and storage efficiency across Gaussian modeling, nearest neighbor retrieval, KV cache compression, and activation quantization tasks.

Bit Budget OptimizationCoordinate PreservationQuantization

This study addresses the suboptimality arising from the decoupling of pruning and quantization in SVD-based compression by proposing a unified co-optimization framework. Methodologically, it introduces a differentiable component-wise bit-width learning mechanism that automatically assigns zero bits to low-importance components, thereby achieving precise pruning. By integrating SVD decomposition with a joint optimization algorithm, the framework enables end-to-end synergy between pruning and quantization. Experimental results demonstrate that under extreme 1.61-bit compression, the proposed approach significantly outperforms two-stage baseline methods. This work establishes a new paradigm for highly efficient compression of large models.

co-optimizationLLM compressionpruning

Hot Scholars

KG

Kun Gai

Senior Director & Researcher, Alibaba Group
Machine LearningComputational Advertising
JH

Jong Hwan Ko

SungKyunKwan Univ. (SKKU)
Deep learning acceleratorImage/audio processingVLSI/IoT systems design
MM

Michele Magno

ETH Zurich
Wireless sensor networksSmart Sensors and Internet of ThingsWake up RadioPower management
YA

Yang Ai

Associate Researcher, University of Science and Technology of China
Speech SynthesisSpeech EnhancementSpeech CodingDeep Learning
HC

Hangting Chen

Tencent Hunyuan
signal processingspeech separationDCASE