discrete latent reasoning

Designs and implements tokenization schemes and models that convert continuous internal states or feature vectors into sequences of discrete latent tokens, and builds or analyzes models that operate over those tokens (for example autoregressive sequence models) to represent, compress, and manipulate intermediate reasoning processes. This work includes discrete latent modeling, latent tokenization, token-based compression of chain-of-thought, and evaluation/decoding methods that preserve stepwise alignment and enable inspection and interpretability.

discretelatentreasoning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.46
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the instability and lack of interpretability in continuous latent-space reasoning, which arises from the mismatch between continuous internal states and discrete symbolic supervision signals. To resolve this, the authors propose Discrete Latent Reasoning (DLR), a novel approach that transforms continuous latent representations into interpretable discrete tokens for the first time. DLR constructs a discrete latent vocabulary by integrating text-to-image rendering, visual feature extraction, and clustering, and unifies the autoregressive modeling of natural language and latent tokens. Inspired by render-and-compress principles, the method achieves state-of-the-art performance across two model families—Qwen3-VL and LLaMA-3—and five reasoning benchmarks, compressing reasoning sequences by up to 20× while preserving semantic fidelity and enabling human-interpretable reasoning trajectories.

continuous latent reasoningdiscrete symbolic supervisioninterpretable reasoning

Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation

Mar 20, 2025
YW
Yuqing Wang
🏛️ University of Hong Kong | ByteDance | Ecole Polytechnique | Peking University

Autoregressive visual generation faces an inherent tension between discrete and continuous token representations: discrete tokens enable simple modeling but suffer from reconstruction distortion and unstable tokenizer training, whereas continuous tokens preserve fidelity at the cost of complex probabilistic modeling. This paper proposes TokenBridge, the first framework to synergistically integrate the advantages of both paradigms. Its core innovation is a decoupled quantization mechanism: post-training, dimension-wise quantization losslessly maps continuous features to discrete tokens—bypassing end-to-end tokenizer optimization—and a lightweight autoregressive prediction head performs standard classification over an ultra-large discrete token space. Experiments demonstrate that TokenBridge achieves reconstruction and generation quality on par with continuous-token baselines, while substantially simplifying training and inference: it requires only cross-entropy loss for efficient optimization, eliminating specialized distribution modeling and tokenizer fine-tuning.

Addresses information loss and instability in discrete token modeling.Bridges continuous and discrete token representations in autoregressive visual generation.Simplifies complex distribution modeling in continuous token approaches.

The Foundations of Tokenization: Statistical and Computational Concerns

Jul 16, 2024
JL
Juan Luis Gastaldi
🏛️ ETH Zürich | City University of New York

This paper addresses the lack of theoretical foundations for tokenization in natural language processing (NLP), systematically investigating its impact on the statistical estimation consistency of language models. While prior work relies predominantly on empirical analysis, we introduce the first unified formal framework grounded in the category of random mappings to rigorously characterize the modeling essence of tokenizers. Our key contributions are: (1) necessary and sufficient conditions for tokenizers to preserve statistical estimation consistency; (2) a four-dimensional theoretical analysis framework—covering inconsistency, ambiguity, finiteness, and sequentiality; and (3) principled, verifiable tokenizer design criteria derived from the integration of category theory, statistical learning theory, and formal language theory. This work establishes the first rigorous mathematical foundation for representation reliability in neural language modeling.

Addressing statistical and computational concerns in tokenizer designAnalyzing tokenization's impact on language model consistencyUnderstanding theoretical foundations of tokenization in NLP

Where is the signal in tokenization space?

Aug 16, 2024
RL
Renato Lui Geh
🏛️ University of California, Los Angeles

This work challenges the conventional assumption in large language models (LLMs) that the probability of a text string equals the probability of its canonical tokenization, revealing that non-canonical tokenizations of the same string encode underutilized semantic and structural signals. We first prove that, under autoregressive LLMs, both finding the most probable tokenization and computing marginal probabilities across all tokenizations are NP-hard. To address this, we propose an efficient approximation algorithm based on dynamic programming with aggressive pruning, compatible with diverse architectures including Transformers and State Space Models (SSMs). Empirically, aggregating marginal probabilities over non-canonical tokenizations—without modifying model parameters or training—yields consistent performance gains across multiple LLM evaluation benchmarks (e.g., LM Evaluation Harness and HELM subsets). These results demonstrate that the tokenization space harbors exploitable latent probabilistic structure, offering a novel, architecture-agnostic avenue for improving LLM inference.

Aggregating non-canonical tokenization probabilities improves LLM benchmark performanceFinding the most likely tokenization for autoregressive LLMs is computationally hardMarginal probability over all tokenizations often matches canonical probability

Toward a Theory of Tokenization in LLMs

Apr 12, 2024
NR
Nived Rajaraman
🏛️ University of California, Berkeley

Why do large language models (LLMs) require tokenization, and why does character-level modeling lead to performance degradation in Transformers? Method: The authors construct a *k*-order Markov data source and rigorously analyze the cross-entropy of Transformers under character-level versus token-level modeling, grounding the analysis in information-theoretic modeling capacity. They establish a provable relationship between tokenization strategies and the accuracy of sequence probability estimation. Contribution/Results: Theoretically, without tokenization, Transformers collapse to modeling only unigram character distributions, failing to capture higher-order dependencies; with appropriate tokenization, learning single-step token predictions suffices to near-optimally model the source distribution. Empirically, tokenization significantly reduces cross-entropy on high-order Markov sources. This work provides the first rigorous information-theoretic and probabilistic justification that tokenization is a necessary condition for overcoming the fundamental limitations of character-level Transformer modeling.

Compare transformers' learning with and without tokenization on simple dataJustify tokenization use by analyzing cross-entropy loss in transformersStudy tokenization's role in transformer performance on Markov processes

Latest Papers

What's happening recently
View more

This work addresses the high computational and memory costs incurred by large language models when processing long prompts, stemming from the quadratic complexity of self-attention. While existing compression methods operate solely in token space and overlook redundancy in the embedding space, this paper introduces K-Token Merging—a novel framework that, for the first time, merges every K consecutive tokens into a single embedding within the latent embedding space via a lightweight encoder. The compressed representation is then processed by a LoRA-finetuned large language model, while generation still employs the original vocabulary. By transcending conventional token-space compression, the method achieves highly efficient input-length reduction with minimal performance degradation. It establishes a Pareto frontier between compression ratio and task performance on Textualized Tree, Amazon Reviews, and CommitPackFT benchmarks, attaining up to 75% compression with negligible loss in accuracy.

computational efficiencyinput length reductionlarge language models

This work addresses the unclear trade-offs among compression efficiency, structural inductive bias, and cross-domain robustness in large language model tokenizers. Viewing tokenization through an information-theoretic lens as structured compression, the authors propose a variant of Byte Pair Encoding (BPE) integrated with principles from compressed sensing and introduce metrics such as channel capacity utilization. They systematically analyze how vocabulary size and training data volume influence text entropy distribution and contextual predictability. Experimental results reveal that while increasing training data enhances token diversity, it simultaneously strengthens contextual predictability. The proposed framework effectively quantifies tokenizer performance, offering both theoretical grounding and practical guidance for designing general-purpose, compression-oriented tokenization strategies and downstream modeling.

compression efficiencydomain shiftlarge language models

本文探讨了语言模型中分词策略对模型训练的影响,通过实验表明输出分词决定了模型的学习难度与内部表示,并指出当前研究对此重视不足。

Autoregressive ModelsModel PerformanceNumeric Reasoning

This study addresses the challenge of aligning generated content with user intent in large language and vision-language models by presenting a systematic review of decoding methods during the inference phase. The research categorizes emerging approaches into three paradigms: token-level guidance, sequence-level generation, and parallel acceleration, while establishing a dedicated resource repository. As the first comprehensive survey of decoding strategy evolution, this work elucidates the critical role of these methods in enhancing both generation efficiency and alignment quality. Furthermore, it outlines future research directions, providing essential theoretical foundations and practical references for optimizing model inference performance.

Decoding MethodsInference-time ControlLarge Language Models

本文提出抽象令牌课程(ATC),一种无需直接监督即可激发有效连续中间表示的框架,解决了大型语言模型中链式思维技术需要丰富任务特定数据的问题。

Abstract Token CurriculumChain-of-ThoughtLarge Language Models

Hot Scholars

ZY

Zitong Yu

U.S. Food and Drug Administration
Medical imagingDeep learningMachine learningImage reconstruction
TT

Tao Tan

FCA MPU
Medical Imaging AI
JH

Jie He

Georgia Institute of Technology
Climate Science
RS

Rui Shao

Professor, Harbin Institute of Technology (Shenzhen)
Computer VisionMultimodal LLMEmbodied AI
PM

Pablo M. Olmos

University Carlos III de Madrid
Machine LearningProbabilistic MethodsMachine Learning for Healthcare