joint vq-ar training

Design and implement training procedures that jointly optimize a vector‑quantized tokenizer/codebook (e.g., VQ‑VAE style) together with an autoregressive generator, including end‑to‑end and tokenizer–generator co‑training strategies. This includes building EMA‑based codebook update rules, mechanisms to detect and restart dead codes, and training schedules or regularizers that maintain full codebook utilization, prevent codebook collapse, and align tokenizer capacity with autoregressive convergence and generator capacity.

jointvq-artraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization

Sep 12, 2025
YC
Yifan Chang
🏛️ CASIA | UCAS | Luoyang Institute for Robot and Intelligent Equipment | Meituan | GigaAI

Vector quantization (VQ) training suffers from poor reconstruction quality and low codebook utilization due to biased straight-through estimator (STE) gradients, gradient sparsity, and delayed codebook updates. To address these issues, this paper proposes VQBridge—a scalable end-to-end trainable framework based on a “compress–process–restore” architecture. VQBridge introduces a novel projection module grounded in functional mapping, refines the STE for more accurate gradient approximation, and integrates a learnable annealing strategy to jointly optimize quantization and reconstruction. This constitutes a new framework for flexible vector quantization (FVQ). Notably, VQBridge achieves 100% codebook utilization for the first time, supports ultra-large codebooks (up to 262K entries), and is compatible with diverse VQ variants. When integrated into LlamaGen, it attains state-of-the-art reconstruction quality, outperforming VAR and DiT by 0.5 and 0.2 in rFID, respectively; moreover, performance improves consistently with increasing codebook size.

Addresses unstable vector quantization training with straight-through estimation biasEnables scalable training for 100% codebook utilization across configurationsSolves low codebook usage and suboptimal reconstruction in VQ networks

Enhancing Vector Quantization with Distributional Matching: A Theoretical and Empirical Study

Jun 18, 2025
XF
Xianghong Fang
🏛️ University of Toronto | The Hong Kong University of Science and Technology | Lehigh University | Southern University of Science and Technology | Boston College

Existing vector quantization (VQ) methods suffer from two key challenges: training instability—caused by gradient mismatch due to the straight-through estimator (STE)—and codebook collapse, stemming from poor codevector utilization. Both issues arise fundamentally from misalignment between the input feature distribution and the codebook distribution. This paper proposes a Wasserstein-distance-based distribution alignment optimization framework—the first to incorporate the Wasserstein distance directly into the VQ objective—to jointly mitigate gradient mismatch and underutilization from a distribution-matching perspective. We provide theoretical guarantees on convergence and quantization error bounds, and design a differentiable distribution alignment mechanism alongside an improved STE. Experiments demonstrate near-perfect codebook utilization (~100%), significantly reduced quantization error, and consistent improvements in reconstruction quality and training stability across multiple autoregressive modeling tasks.

Addresses training instability in vector quantization methodsReduces quantization error via feature-code distribution alignmentResolves codebook collapse by improving code vector utilization

This work identifies and addresses the "entropy cliff" phenomenon in conventional discrete visual autoregressive models, where a fixed codebook causes a sharp drop in conditional entropy toward the end of the sequence, reducing generation to mere memorization and limiting reconstruction fidelity. To overcome this fundamental limitation, the authors propose Variable Codebook Quantization (VCQ), which monotonically increases codebook capacity along the sequence—from a minimum size \(K_{\text{min}} = 2\) to a maximum \(K_{\text{max}}\)—within a standard autoregressive Transformer architecture. Notably, VCQ requires no modifications to the loss function, model parameters, or training protocol, yet induces a coarse-to-fine semantic hierarchy. On ImageNet at 256×256 resolution, the base model reduces gFID from 27.98 to 14.80, with an extended variant achieving 1.71; furthermore, a linear probe using only the first 10 tokens attains 43.8% top-1 accuracy, surpassing the information-theoretic bottleneck imposed by fixed codebooks.

autoregressive visual generationcodebook sizeconditional entropy

This work addresses the significant challenges in post-training quantization (PTQ) of autoregressive vision generation (ARVG) models, namely channel-wise outliers, highly dynamic token-level activations, and mismatched sample distributions. It presents the first systematic analysis of these issues and introduces PTQ4ARVG, a training-free quantization framework comprising three key components: Gain-Projected Scaling (GPS), Static Token-Wise Quantization (STWQ), and Distribution-Guided Calibration (DGC), integrated with Taylor expansion and entropy-driven sampling strategies. Without requiring any fine-tuning, PTQ4ARVG successfully compresses diverse ARVG models to 8-bit and even 6-bit precision, substantially reducing model size and inference latency while preserving generation quality, thereby achieving notable improvements in quantization accuracy and generalization capability.

Activation DistributionAutoRegressive Visual GenerationOutliers

Optimal and Near-Optimal Adaptive Vector Quantization

Feb 05, 2024
RB
Ran Ben-Basat
🏛️ UCL | VMware Research | Harvard

Adaptive vector quantization (AVQ) is essential for compressing gradients, weights, activations, and datasets in machine learning, yet existing methods suffer from prohibitive time and memory complexity. Method: We propose the first algorithms for AVQ—both strictly optimal and highly efficient near-optimal—overcoming these bottlenecks. Our optimal algorithm employs a progressive dynamic programming framework with greedy pruning and error-bounded divide-and-conquer. For large-scale inputs, we introduce a super-fast near-optimal variant leveraging structural approximations and rigorous error analysis. Contribution/Results: The optimal algorithm guarantees theoretical precision, while the near-optimal variant achieves controllable distortion with 10–100× speedup and significantly reduced memory footprint. Both support seamless end-to-end integration into modern ML systems, enabling practical AVQ deployment across training and inference pipelines.

Enable efficient AVQ for large-scale inputsOptimize adaptive vector quantization for machine learningReduce runtime and memory of optimal quantization methods

Latest Papers

What's happening recently
View more

This work addresses the performance limitations of conventional two-stage training in visual generative models, where the tokenizer and generator develop misaligned modeling preferences. To overcome this, the authors propose GEAR, a novel approach featuring a hard-soft dual-branch mechanism that enables end-to-end joint training of vector-quantized tokenizers and autoregressive generators through representation alignment. While preserving the autoregressive property, GEAR steers the tokenizer toward learning index distributions that are easier for the generator to predict, effectively shifting the burden of semantic alignment to the generator. The method is compatible with various quantizers—including VQVAE, LFQ, and IBQ—and achieves up to a 10× faster convergence in gFID on ImageNet. It also substantially improves patch-level and spatial consistency and successfully generalizes to text-to-image generation tasks.

autoregressive generationend-to-end trainingimage synthesis

This work addresses the instability and codebook collapse in vector quantization caused by distributional mismatch between features and the codebook by introducing distribution matching as a central principle and proposing a unified theoretical framework. The approach explicitly aligns the two distributions using either the Wasserstein distance—admitting a closed-form solution under Gaussian approximation—or a non-parametric maximum mean discrepancy (MMD). This alignment significantly enhances codebook utilization and stabilizes training. Experimental results demonstrate that the proposed method substantially outperforms existing approaches on visual tokenization benchmarks, exhibiting strong effectiveness, robustness, and efficient codebook usage.

codebook collapsedistributional mismatchfeature discretization

This work addresses codebook collapse in large-scale vector quantization, a phenomenon characterized by unassigned codewords and increased quantization error, and identifies encoder drift as the primary underlying cause. To mitigate this issue, the authors propose Non-stationary-aware Vector Quantization (NSVQ), a novel training strategy that employs a non-stationary embedding loss to guide the codebook in tracking early-stage encoder drift. NSVQ integrates dynamic codebook replacement and phased encoder freezing to first enable joint optimization and subsequently stabilize the architecture, followed by adversarial fine-tuning to disrupt the feedback loop of quantization error. Evaluated on ImageNet-1k at 128×128 resolution, NSVQ reduces the reconstruction Fréchet Inception Distance (rFID) from 2.39 to 2.10, achieves 100% codebook utilization, and substantially enhances the generation quality of downstream latent diffusion models.

codebook collapseencoder driftgenerative modeling

Training state-of-the-art vision vector quantization (VQ) models is computationally expensive, hindering innovation in resource-constrained settings. This work proposes a plug-and-play framework that enables direct integration of novel VQ modules into frozen, pre-trained visual tokenizers without end-to-end retraining. Feature prior alignment is achieved through a lightweight decoder adaptation strategy involving only five epochs of fine-tuning on ImageNet-1k. The method achieves, for the first time, efficient VQ module replacement without retraining, attaining near state-of-the-art reconstruction fidelity on industrial-scale models such as VAR while reducing training costs by 95%. This dramatic reduction in computational requirements significantly lowers the barrier to entry and advances the democratization of VQ technologies.

Computational CostQuantization ModuleResource Constraints

Hot Scholars

AF

Aasa Feragen

Professor, DTU Compute
Machine learningmedical imaginggeometric modelling
ML

Miguel López-Pérez

Postdoc at Universitat Politècnica de València
Gaussian ProcessesDigital PathologyCrowdsourcingWeakly Supervised Learning
SH

Søren Hauberg

Cognitive Systems, DTU Compute, Technical University of Denmark
Machine LearningComputer VisionGeometric Statistics
XG

Xiang-Gen Xia

Department of Electrical and Computer Engineering, University of Delaware, Newark, DE 19716, USA
signal processingdigital communicationsradar signal processing
ZY

Zitong Yu

U.S. Food and Drug Administration
Medical imagingDeep learningMachine learningImage reconstruction