Score
Designs and implements tokenization schemes and models that convert continuous internal states or feature vectors into sequences of discrete latent tokens, and builds or analyzes models that operate over those tokens (for example autoregressive sequence models) to represent, compress, and manipulate intermediate reasoning processes. This work includes discrete latent modeling, latent tokenization, token-based compression of chain-of-thought, and evaluation/decoding methods that preserve stepwise alignment and enable inspection and interpretability.
Discrete tokenizers lack a systematic, cross-task survey. Method: We propose the first unified analytical framework covering generation, understanding, recommendation, and information retrieval; introduce a hierarchical decomposition paradigm for tokenizer submodules; establish a cross-task taxonomy; and conduct a horizontal comparison of representative approaches—including VQ-VAE, SoundStream, K-means tokenization, semantic hashing, and cross-modal alignment—through the lenses of information theory, representation learning, and structured modeling. Contribution/Results: We identify three core challenges: semantic alignment, cross-modal generalization, and the efficiency–accuracy trade-off. Furthermore, we deliver a reusable evaluation dimension matrix and an open challenge map, providing both theoretical foundations and practical guidelines for designing next-generation tokenizers that are robust, interpretable, and cross-modal.
This work addresses the instability and lack of interpretability in continuous latent-space reasoning, which arises from the mismatch between continuous internal states and discrete symbolic supervision signals. To resolve this, the authors propose Discrete Latent Reasoning (DLR), a novel approach that transforms continuous latent representations into interpretable discrete tokens for the first time. DLR constructs a discrete latent vocabulary by integrating text-to-image rendering, visual feature extraction, and clustering, and unifies the autoregressive modeling of natural language and latent tokens. Inspired by render-and-compress principles, the method achieves state-of-the-art performance across two model families—Qwen3-VL and LLaMA-3—and five reasoning benchmarks, compressing reasoning sequences by up to 20× while preserving semantic fidelity and enabling human-interpretable reasoning trajectories.
Autoregressive visual generation faces an inherent tension between discrete and continuous token representations: discrete tokens enable simple modeling but suffer from reconstruction distortion and unstable tokenizer training, whereas continuous tokens preserve fidelity at the cost of complex probabilistic modeling. This paper proposes TokenBridge, the first framework to synergistically integrate the advantages of both paradigms. Its core innovation is a decoupled quantization mechanism: post-training, dimension-wise quantization losslessly maps continuous features to discrete tokens—bypassing end-to-end tokenizer optimization—and a lightweight autoregressive prediction head performs standard classification over an ultra-large discrete token space. Experiments demonstrate that TokenBridge achieves reconstruction and generation quality on par with continuous-token baselines, while substantially simplifying training and inference: it requires only cross-entropy loss for efficient optimization, eliminating specialized distribution modeling and tokenizer fine-tuning.
This paper addresses the lack of theoretical foundations for tokenization in natural language processing (NLP), systematically investigating its impact on the statistical estimation consistency of language models. While prior work relies predominantly on empirical analysis, we introduce the first unified formal framework grounded in the category of random mappings to rigorously characterize the modeling essence of tokenizers. Our key contributions are: (1) necessary and sufficient conditions for tokenizers to preserve statistical estimation consistency; (2) a four-dimensional theoretical analysis framework—covering inconsistency, ambiguity, finiteness, and sequentiality; and (3) principled, verifiable tokenizer design criteria derived from the integration of category theory, statistical learning theory, and formal language theory. This work establishes the first rigorous mathematical foundation for representation reliability in neural language modeling.
This work challenges the conventional assumption in large language models (LLMs) that the probability of a text string equals the probability of its canonical tokenization, revealing that non-canonical tokenizations of the same string encode underutilized semantic and structural signals. We first prove that, under autoregressive LLMs, both finding the most probable tokenization and computing marginal probabilities across all tokenizations are NP-hard. To address this, we propose an efficient approximation algorithm based on dynamic programming with aggressive pruning, compatible with diverse architectures including Transformers and State Space Models (SSMs). Empirically, aggregating marginal probabilities over non-canonical tokenizations—without modifying model parameters or training—yields consistent performance gains across multiple LLM evaluation benchmarks (e.g., LM Evaluation Harness and HELM subsets). These results demonstrate that the tokenization space harbors exploitable latent probabilistic structure, offering a novel, architecture-agnostic avenue for improving LLM inference.
Why do large language models (LLMs) require tokenization, and why does character-level modeling lead to performance degradation in Transformers? Method: The authors construct a *k*-order Markov data source and rigorously analyze the cross-entropy of Transformers under character-level versus token-level modeling, grounding the analysis in information-theoretic modeling capacity. They establish a provable relationship between tokenization strategies and the accuracy of sequence probability estimation. Contribution/Results: Theoretically, without tokenization, Transformers collapse to modeling only unigram character distributions, failing to capture higher-order dependencies; with appropriate tokenization, learning single-step token predictions suffices to near-optimally model the source distribution. Empirically, tokenization significantly reduces cross-entropy on high-order Markov sources. This work provides the first rigorous information-theoretic and probabilistic justification that tokenization is a necessary condition for overcoming the fundamental limitations of character-level Transformer modeling.
This work addresses the high computational and memory costs incurred by large language models when processing long prompts, stemming from the quadratic complexity of self-attention. While existing compression methods operate solely in token space and overlook redundancy in the embedding space, this paper introduces K-Token Merging—a novel framework that, for the first time, merges every K consecutive tokens into a single embedding within the latent embedding space via a lightweight encoder. The compressed representation is then processed by a LoRA-finetuned large language model, while generation still employs the original vocabulary. By transcending conventional token-space compression, the method achieves highly efficient input-length reduction with minimal performance degradation. It establishes a Pareto frontier between compression ratio and task performance on Textualized Tree, Amazon Reviews, and CommitPackFT benchmarks, attaining up to 75% compression with negligible loss in accuracy.
This work addresses the unclear trade-offs among compression efficiency, structural inductive bias, and cross-domain robustness in large language model tokenizers. Viewing tokenization through an information-theoretic lens as structured compression, the authors propose a variant of Byte Pair Encoding (BPE) integrated with principles from compressed sensing and introduce metrics such as channel capacity utilization. They systematically analyze how vocabulary size and training data volume influence text entropy distribution and contextual predictability. Experimental results reveal that while increasing training data enhances token diversity, it simultaneously strengthens contextual predictability. The proposed framework effectively quantifies tokenizer performance, offering both theoretical grounding and practical guidance for designing general-purpose, compression-oriented tokenization strategies and downstream modeling.
本文探讨了语言模型中分词策略对模型训练的影响,通过实验表明输出分词决定了模型的学习难度与内部表示,并指出当前研究对此重视不足。
This study addresses the challenge of aligning generated content with user intent in large language and vision-language models by presenting a systematic review of decoding methods during the inference phase. The research categorizes emerging approaches into three paradigms: token-level guidance, sequence-level generation, and parallel acceleration, while establishing a dedicated resource repository. As the first comprehensive survey of decoding strategy evolution, this work elucidates the critical role of these methods in enhancing both generation efficiency and alignment quality. Furthermore, it outlines future research directions, providing essential theoretical foundations and practical references for optimizing model inference performance.
本文提出抽象令牌课程(ATC),一种无需直接监督即可激发有效连续中间表示的框架,解决了大型语言模型中链式思维技术需要丰富任务特定数据的问题。