Highly Compressed Tokenizer Can Generate Without Training

📅 2025-06-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Conventional image editing and generation rely on large, pre-trained generative models, incurring high computational and data requirements. Method: We propose a training-free paradigm based on a one-dimensional vector-quantized tokenizer (1D VQ-Tokenizer) with only 32 discrete tokens, achieving extreme compression (~1024×) while preserving rich semantic structure in the latent space. Fine-grained editing is performed via token-level heuristic operations (e.g., copy, replace), and end-to-end generation is realized through test-time gradient optimization with plug-and-play losses—reconstruction and CLIP-guided similarity. Contribution/Results: We empirically demonstrate, for the first time, that this 1D latent space supports strong semantic editability and generative capability without model training. It enables zero-shot inpainting and text-guided editing. Experiments show competitive diversity and photorealism compared to supervised methods, while drastically reducing computational cost and data dependency.

Technology Category

Computer Vision: Diffusion Models for VisionMachine Learning: Deep Generative Models & AutoencodersNatural Language Processing: Generation

Application Category

Economics, Online Markets and Human Computation: Economic ramifications for generative AI infrastructure and applicationsWeb Mining and Content Analysis: Large pretrained models with web dataSearch and Retrieval-Augmented AI: Large language models for search
📝 Abstract
Commonly used image tokenizers produce a 2D grid of spatially arranged tokens. In contrast, so-called 1D image tokenizers represent images as highly compressed one-dimensional sequences of as few as 32 discrete tokens. We find that the high degree of compression achieved by a 1D tokenizer with vector quantization enables image editing and generative capabilities through heuristic manipulation of tokens, demonstrating that even very crude manipulations -- such as copying and replacing tokens between latent representations of images -- enable fine-grained image editing by transferring appearance and semantic attributes. Motivated by the expressivity of the 1D tokenizer's latent space, we construct an image generation pipeline leveraging gradient-based test-time optimization of tokens with plug-and-play loss functions such as reconstruction or CLIP similarity. Our approach is demonstrated for inpainting and text-guided image editing use cases, and can generate diverse and realistic samples without requiring training of any generative model.
Problem

Research questions and friction points this paper is trying to address.

Explores 1D image tokenizers for high compression
Enables image editing via heuristic token manipulation
Generates images without training using optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

1D tokenizer enables image editing via token manipulation
Gradient-based optimization with plug-and-play loss functions
Generates images without training any generative model