🤖 AI Summary
Training state-of-the-art vision vector quantization (VQ) models is computationally expensive, hindering innovation in resource-constrained settings. This work proposes a plug-and-play framework that enables direct integration of novel VQ modules into frozen, pre-trained visual tokenizers without end-to-end retraining. Feature prior alignment is achieved through a lightweight decoder adaptation strategy involving only five epochs of fine-tuning on ImageNet-1k. The method achieves, for the first time, efficient VQ module replacement without retraining, attaining near state-of-the-art reconstruction fidelity on industrial-scale models such as VAR while reducing training costs by 95%. This dramatic reduction in computational requirements significantly lowers the barrier to entry and advances the democratization of VQ technologies.
📝 Abstract
Vector Quantization (VQ) underpins modern discrete visual tokenization. However, training quantization modules for state-of-the-art VQ-based models requires significant computational resources which, in practice, all but prevents the development of novel, cutting-edge VQ techniques under resource constraints. To address this limitation, we propose {\bf VQ-Transplant}, a simple framework that enables plug-and-play integration of new VQ modules into frozen, pre-trained tokenizers by replacing their native VQ modules. Crucially, the proposed transplantation process preserves all encoder-decoder parameters, obviating the need for costly end-to-end retraining when modifying the quantization method. To mitigate decoder-quantization mismatch, we introduce a lightweight decoder adaptation strategy (trained for only 5 epochs on ImageNet-1k) to align feature priors with the new quantization space. In our empirical evaluation, we find that VQ-Transplant allows obtaining near state-of-the-art reconstruction fidelity for industry-level models like VAR while reducing the training cost by 95\%. VQ-Transplant democratizes quantization research by enabling resource-efficient integration of novel VQ techniques while matching industry-level reconstruction performance.