Information-Theoretic Analysis of Next-Token Prediction under Markovian Data

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear generalization mechanisms of next-token prediction under Markov data. To this end, it constructs an information-theoretic framework that decouples algorithmic effects from temporal dependencies. By introducing rate-distortion theory to accommodate continuous hypothesis spaces and integrating the Donsker-Varadhan variational representation, McDiarmid’s inequality, and noisy low-dimensional compression techniques, the authors derive generalization bounds for both cross-entropy and margin-based predictions. The work elucidates how mixing properties influence generalization and reveals that extended contexts widen the generalization gap. Furthermore, experiments on the ETTh2 dataset empirically validate 24 hours as an effective memory horizon.
📝 Abstract
We develop an information-theoretic framework for generalization in next-token prediction under temporally dependent data. We consider independent trajectories generated by finite-memory Markov processes and distinguish algorithmic dependence, quantified by mutual information, from temporal dependence, characterized by mixing. For cross-entropy loss, we derive an expected generalization bound using the Donsker--Varadhan variational representation and a McDiarmid-type concentration inequality for Markov chains. A refinement captures the joint effect of context length and temporal mixing through the mixing properties of the history-state process. We then extend the bound through a rate--distortion formulation, replacing mutual information with the minimum information rate required to represent the learned model within a prescribed distortion in the generalization gap, yielding informative guarantees for deterministic algorithms over continuous hypothesis spaces. For margin-based prediction, we derive explicit bounds for linear and self-attention next-token predictors via noisy low-dimensional compression, revealing the roles of context length, model complexity, sample size, margin, and temporal mixing. Experiments on TinyStories and ETTh2 show that longer contexts can reduce both training and test losses, but typically reduce training loss more, enlarging the generalization gap. A complementary ETTh2 analysis identifies an effective predictive-memory scale near 24 hours, with no statistically supported improvement beyond this scale, offering a plausible explanation for test-performance saturation at larger contexts.
Problem

Research questions and friction points this paper is trying to address.

next-token prediction
generalization bound
Markovian data
information theory
context length
Innovation

Methods, ideas, or system contributions that make the work stand out.

Information-Theoretic Framework
Next-Token Prediction
Markovian Data
Rate-Distortion
Generalization Bound
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Masoud Kavian
Masoud Kavian
Huawei Paris Research Center
Information TheoryLearning Theory
A
Abdellatif Zaidi
Mathematical and Algorithmic Science Laboratory, Huawei Paris Research Center, 92100 Boulogne-Billancourt, France; and Laboratoire d’Informatique Gaspard Monge, Université Gustave Eiffel, 77420 Champs-sur-Marne, France
Milad Sefidgaran
Milad Sefidgaran
Senior ML Researcher
Machine LearningDeep LearningInformation Theory