🤖 AI Summary
Conventional image editing and generation rely on large, pre-trained generative models, incurring high computational and data requirements. Method: We propose a training-free paradigm based on a one-dimensional vector-quantized tokenizer (1D VQ-Tokenizer) with only 32 discrete tokens, achieving extreme compression (~1024×) while preserving rich semantic structure in the latent space. Fine-grained editing is performed via token-level heuristic operations (e.g., copy, replace), and end-to-end generation is realized through test-time gradient optimization with plug-and-play losses—reconstruction and CLIP-guided similarity. Contribution/Results: We empirically demonstrate, for the first time, that this 1D latent space supports strong semantic editability and generative capability without model training. It enables zero-shot inpainting and text-guided editing. Experiments show competitive diversity and photorealism compared to supervised methods, while drastically reducing computational cost and data dependency.
📝 Abstract
Commonly used image tokenizers produce a 2D grid of spatially arranged tokens. In contrast, so-called 1D image tokenizers represent images as highly compressed one-dimensional sequences of as few as 32 discrete tokens. We find that the high degree of compression achieved by a 1D tokenizer with vector quantization enables image editing and generative capabilities through heuristic manipulation of tokens, demonstrating that even very crude manipulations -- such as copying and replacing tokens between latent representations of images -- enable fine-grained image editing by transferring appearance and semantic attributes. Motivated by the expressivity of the 1D tokenizer's latent space, we construct an image generation pipeline leveraging gradient-based test-time optimization of tokens with plug-and-play loss functions such as reconstruction or CLIP similarity. Our approach is demonstrated for inpainting and text-guided image editing use cases, and can generate diverse and realistic samples without requiring training of any generative model.