Alternatives To Next Token Prediction In Text Generation - A Survey

📅 2025-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Next-token prediction (NTP)-based language models suffer from fundamental limitations: weak long-horizon planning, severe error accumulation, and low computational efficiency. To address these, this paper systematically surveys alternative paradigms to NTP and proposes, for the first time, a unified five-dimensional taxonomy: multi-token prediction, plan-then-generate, latent-space reasoning, continuous-generation methods, and non-Transformer architectures. By integrating techniques—including multi-step forecasting, hierarchical planning, continuous latent-space modeling, diffusion/flow-matching, energy-based optimization, and novel neural structures—the work characterizes performance boundaries and synergistic potential across approaches. Crucially, it establishes the first comprehensive, dimensionally explicit methodology for NTP alternatives. This framework provides both theoretical foundations and concrete technical pathways toward developing efficient, controllable, and high-fidelity text generation models.

Technology Category

Natural Language Processing: GenerationMachine Learning: Large Multimodal Models (LMMs)Planning, Routing, and Scheduling: Planning with Language Models

Application Category

Search and Retrieval-Augmented AI: Large language models for searchSocial Networks and Social Media: Generative AI / large language models and their impact on social systemsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
The paradigm of Next Token Prediction (NTP) has driven the unprecedented success of Large Language Models (LLMs), but is also the source of their most persistent weaknesses such as poor long-term planning, error accumulation, and computational inefficiency. Acknowledging the growing interest in exploring alternatives to NTP, the survey describes the emerging ecosystem of alternatives to NTP. We categorise these approaches into five main families: (1) Multi-Token Prediction, which targets a block of future tokens instead of a single one; (2) Plan-then-Generate, where a global, high-level plan is created upfront to guide token-level decoding; (3) Latent Reasoning, which shifts the autoregressive process itself into a continuous latent space; (4) Continuous Generation Approaches, which replace sequential generation with iterative, parallel refinement through diffusion, flow matching, or energy-based methods; and (5) Non-Transformer Architectures, which sidestep NTP through their inherent model structure. By synthesizing insights across these methods, this survey offers a taxonomy to guide research into models that address the known limitations of token-level generation to develop new transformative models for natural language processing.
Problem

Research questions and friction points this paper is trying to address.

Surveying alternatives to Next Token Prediction for text generation
Addressing poor long-term planning and error accumulation in LLMs
Categorizing five families of methods to overcome NTP limitations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-token prediction for future token blocks
Plan-then-generate with upfront global planning
Latent reasoning in continuous autoregressive space
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Charlie Wyatt
School of Computer Science and Engineering, UNSW Sydney, Sydney, Australia
A
Aditya Joshi
School of Computer Science and Engineering, UNSW Sydney, Sydney, Australia
Flora Salim
Flora Salim
Professor, CSE, UNSW
Machine LearningTime SeriesSpatiotemporalUbiCompFoundation Models