LightMTP: Lightweight Latent Multi-Token Prediction

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the substantial parameter overhead and reliance on external model supervision inherent in existing multi-token prediction methods by proposing a lightweight latent multi-token prediction framework. The approach leverages bootstrapped hidden-state encoding, wherein the model's own hidden layers guide the learning of latent representations for future tokens, eliminating the need for additional computation or external supervisory signals. Experimental results demonstrate that this framework achieves parameter-efficient latent encoding with only a 1% increase in parameter count. Furthermore, it preserves general-purpose performance while significantly enhancing planning, programming, and reasoning capabilities, effectively reducing the resource consumption associated with multi-token extended supervision.
📝 Abstract
Next-token prediction (NTP) is the standard pretraining objective for large language models, yet it provides an explicit training signal only for the immediate next token, which can lead models to exploit local patterns instead of capturing longer-range structure and ideas. Multi-token prediction (MTP) addresses this by training models to predict several future tokens. However, existing MTP methods often introduce a large number of new parameters with limited improvements in downstream performance. Latent MTP approaches address this efficiency issue by encoding future tokens into a vector representation. However, these approaches usually rely on external helper models for future token encoding. We propose LightMTP, a lightweight, i.e., parameter-efficient, latent MTP approach that bootstraps the future token representations from the model's own hidden states. Our two LightMTP variants extend supervision to more future tokens without requiring the additional computational overhead of conventional MTP nor the external supervision latent MTP normally relies on. LightMTP adds at most 1% extra parameters, retains better performance on general language modeling benchmarks, and achieves similar gains in planning, coding, and reasoning.
Problem

Research questions and friction points this paper is trying to address.

Next-token prediction
Multi-token prediction
Latent multi-token prediction
Parameter efficiency
External helper models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Token Prediction
Latent Representation
Parameter-Efficient
Self-Bootstrapping
Large Language Models
🔎 Similar Papers
No similar papers found.