Score
Designs, implements, and trains methods that convert visual inputs into compact discrete or subword token sequences and associated embeddings—covering geometry-aware, input-adaptive, and modality-specific tokenizers, per-token weighting, and mechanisms for token embedding design, fusion, and manipulation. Measures and analyzes token distributions, frequencies, probabilities and usage, evaluates token-efficiency and reconstruction fidelity, enforces semantic discriminability, and implements token-level alignment, conditioning, and authentication for downstream attention and conditioning workflows.
Discrete tokenizers lack a systematic, cross-task survey. Method: We propose the first unified analytical framework covering generation, understanding, recommendation, and information retrieval; introduce a hierarchical decomposition paradigm for tokenizer submodules; establish a cross-task taxonomy; and conduct a horizontal comparison of representative approaches—including VQ-VAE, SoundStream, K-means tokenization, semantic hashing, and cross-modal alignment—through the lenses of information theory, representation learning, and structured modeling. Contribution/Results: We identify three core challenges: semantic alignment, cross-modal generalization, and the efficiency–accuracy trade-off. Furthermore, we deliver a reusable evaluation dimension matrix and an open challenge map, providing both theoretical foundations and practical guidelines for designing next-generation tokenizers that are robust, interpretable, and cross-modal.
研究通过构建纯自回归测试平台,分析图像标记器与文本联合建模时的行为,探讨了不同任务下的损失与下游性能的关系,以及图像标记器设计对联合优化的影响。
This study investigates how the geometric structure of token embeddings in large language models (LLMs) influences next-token prediction performance. Method: We propose a systematic empirical measurement framework that integrates geometric metrics—including intrinsic dimensionality, neighborhood overlap, and cosine similarity—and quantitatively correlates them with per-layer cross-entropy loss for the first time. Contribution/Results: We find that high prediction loss consistently corresponds to higher-dimensional embedding spaces, revealing an implicit link between geometric dimensionality and model uncertainty; this correlation vanishes under token shuffling, confirming the decisive role of syntactic and semantic structure in shaping embedding geometry. Furthermore, layer-wise geometric evolution analysis characterizes a dynamic transition from discretized to manifold-like representations across layers. These findings establish embedding geometry as an interpretable proxy for predictive performance, offering a novel perspective for model diagnosis and optimization.
Vision Transformers (ViTs) rely on uniform image patching, yielding tokens with limited semantic meaning and compositional structure. Method: We propose a semantics-driven visual tokenization framework that replaces fixed-size patches with “tangible tokens” derived from instance segmentation masks and “intangible tokens” extracted from scene graphs to encode relationships and actions. We introduce an additive attention mechanism to explicitly model structural and semantic dependencies among tokens, and design an end-to-end vision-language pretraining framework enabling fine-grained alignment between visual token embeddings and textual caption embeddings. Contribution/Results: This work is the first to systematically incorporate interpretable, semantically grounded visual tokens. On COCO image-text retrieval, it improves text-to-image and image-to-text recall by 47% and 44%, respectively. On compositional reasoning benchmarks—ARO and Winoground—it achieves gains of 18% and 10%, significantly enhancing semantic understanding and compositional generalization.
This work addresses the challenge in vision-language models (VLMs) of simultaneously achieving discriminative and generative capabilities for custom concepts, along with insufficient composability. We propose a composable custom token learning framework that jointly optimizes text inversion loss and classification loss using only a few images and textual definitions of parent classes, while enforcing regularization via a low-dimensional attribute embedding subspace. We introduce Generative-Aided Information Retrieval (GAIR), a novel paradigm enabling a single custom token to function effectively and coherently across classification, cross-modal retrieval, and text-to-image generation. On DeepFashion2, our method improves mean reciprocal rank (MRR) for text-to-image retrieval by 7%. It further supports interpretable composite query visualization and dynamic inference-time correction, significantly enhancing generalization and controllability for novel concepts.
To address the performance trade-off between multimodal understanding and generation caused by mismatched visual granularity, this paper proposes TokenFlow—the first unified image tokenizer. Its core innovation is a dual-codebook vector quantization architecture: a semantic codebook captures high-level semantics, while a pixel codebook preserves fine-grained texture; both are decoupled during learning yet jointly aligned via a shared index mechanism. This design enables discrete token-driven, integrated modeling of multimodal understanding and autoregressive image generation. Experiments demonstrate that TokenFlow achieves an average 7.2% improvement over LLaVA-1.5 (13B) on multimodal understanding benchmarks; attains an FID of 0.63 for 384×384 image reconstruction; and scores 0.55 on GenEval for 256×256 autoregressive generation—comparable to SDXL. TokenFlow effectively mitigates granularity conflict and advances unified visual representation learning.
This study investigates whether visual tokens in vision-language models require sustained processing in deep layers. Through analyses of representation entropy, intrinsic dimensionality, and trajectory curvature—combined with depth truncation and deterministic decoding experiments—the authors find that visual representations rapidly converge to a stable state in shallow layers, rendering deeper processing largely inconsequential for final outputs. The work further reveals substantial interchangeability of visual tokens across layers, challenging the prevailing paradigm that multimodal large language models critically depend on deep visual processing. Additional analysis shows that single-token prediction is robust to visual depth truncation, whereas multi-token generation tasks rely on continuous visual input; crucially, deep visual processing primarily shapes the inference trajectory rather than the ultimate result.
This study addresses the degradation of spatial localization caused by visual token pruning and the high computational overhead of resamplers under extreme compression. To this end, we propose Braco, a lightweight encoder that decouples compressibility from learnability objectives. Braco introduces a four-step parameterized encoding pipeline incorporating orthogonal reparameterization, input-agnostic basis coordinate embedding, and spatial residual token pooling, achieving efficient visual compression through basis transform truncation and coordinate reassembly. Experimental results demonstrate that Braco preserves 95.2% localization accuracy at a 64× compression ratio while reducing FLOPs by over 84% and improving inference speed by 36%, thereby establishing an effective balance between extreme compression and model performance.
本文提出了一种新架构Giraffe,通过单个[IMG]标记将隐藏文本表示映射到视觉嵌入空间,以解决图形设计生成中多模态大语言模型生成媒体能力有限的问题。
This work addresses the severe degradation in text legibility and facial details caused by coarse-grained downsampling and quantization in standard tokenizers for discrete autoregressive image generation. To mitigate this, the authors propose a content-aware, region-sensitive reconstruction loss that explicitly optimizes the fidelity of text and face regions during tokenizer training. This approach introduces, for the first time in discrete visual tokenizers, semantic-region-specific supervision signals. Combined with a compact 16k codebook and a 16× downsampling ratio, it enables fine-grained and adaptive reconstruction. Experiments demonstrate that the proposed tokenizer significantly outperforms existing methods in reconstructing text and facial regions, and successfully transfers to the InsightAR generative model, enhancing textual clarity and facial realism while preserving strong general reconstruction performance.
This work proposes SemTok, a semantic-driven one-dimensional tokenizer that addresses the limitations of existing vision tokenizers, which often rely on fixed 2D grids and prioritize pixel-level reconstruction at the expense of compact global semantics. SemTok compresses images into high-level discrete semantic tokens and introduces a masked autoregressive generative framework. Its key innovations include a 2D-to-1D semantic tokenization strategy, a semantic alignment constraint mechanism, and a two-stage generative training paradigm. Experimental results demonstrate that SemTok achieves state-of-the-art performance in image reconstruction, delivering higher fidelity under extremely compact token representations and significantly enhancing downstream generative capabilities.