Group-Wise Optimization for Self-Extensible Codebooks in Vector Quantized Models

📅 2025-10-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
VQ-VAEs suffer from codebook collapse in self-supervised vector reconstruction, and existing approaches—either employing implicit static codebooks or jointly optimizing the entire codebook—constrain representational capacity, degrading reconstruction fidelity. To address this, we propose Grouped Vector Quantization (Group-VQ): the codebook is partitioned into disjoint groups; vectors within each group are jointly optimized, while groups are updated independently, thereby enhancing codebook utilization. Additionally, we introduce a post-training codebook resampling mechanism that dynamically expands the codebook size without requiring retraining. This design achieves joint optimization of codebook efficiency and reconstruction performance while preserving model lightweightness. Experiments across multiple image reconstruction benchmarks demonstrate significant improvements in PSNR and LPIPS, validating Group-VQ’s effectiveness and generalizability.

Technology Category

Machine Learning: Deep Generative Models & AutoencodersComputer Vision: Learning & Optimization for CVSearch and Optimization: Sampling/Simulation-based Search

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphsUser Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendation
📝 Abstract
Vector Quantized Variational Autoencoders (VQ-VAEs) leverage self-supervised learning through reconstruction tasks to represent continuous vectors using the closest vectors in a codebook. However, issues such as codebook collapse persist in the VQ model. To address these issues, existing approaches employ implicit static codebooks or jointly optimize the entire codebook, but these methods constrain the codebook's learning capability, leading to reduced reconstruction quality. In this paper, we propose Group-VQ, which performs group-wise optimization on the codebook. Each group is optimized independently, with joint optimization performed within groups. This approach improves the trade-off between codebook utilization and reconstruction performance. Additionally, we introduce a training-free codebook resampling method, allowing post-training adjustment of the codebook size. In image reconstruction experiments under various settings, Group-VQ demonstrates improved performance on reconstruction metrics. And the post-training codebook sampling method achieves the desired flexibility in adjusting the codebook size.
Problem

Research questions and friction points this paper is trying to address.

Addresses codebook collapse in vector quantized models
Improves trade-off between codebook utilization and reconstruction
Enables post-training adjustment of codebook size flexibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

Group-wise optimization for independent codebook training
Joint optimization performed within each codebook group
Training-free resampling enables post-training codebook adjustment
🔎 Similar Papers
No similar papers found.
H
Hong-Kai Zheng
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics, China
P
Piji Li
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics, China