🤖 AI Summary
This study addresses the unclear generalization mechanisms of next-token prediction under Markov data. To this end, it constructs an information-theoretic framework that decouples algorithmic effects from temporal dependencies. By introducing rate-distortion theory to accommodate continuous hypothesis spaces and integrating the Donsker-Varadhan variational representation, McDiarmid’s inequality, and noisy low-dimensional compression techniques, the authors derive generalization bounds for both cross-entropy and margin-based predictions. The work elucidates how mixing properties influence generalization and reveals that extended contexts widen the generalization gap. Furthermore, experiments on the ETTh2 dataset empirically validate 24 hours as an effective memory horizon.
📝 Abstract
We develop an information-theoretic framework for generalization in next-token prediction under temporally dependent data. We consider independent trajectories generated by finite-memory Markov processes and distinguish algorithmic dependence, quantified by mutual information, from temporal dependence, characterized by mixing. For cross-entropy loss, we derive an expected generalization bound using the Donsker--Varadhan variational representation and a McDiarmid-type concentration inequality for Markov chains. A refinement captures the joint effect of context length and temporal mixing through the mixing properties of the history-state process. We then extend the bound through a rate--distortion formulation, replacing mutual information with the minimum information rate required to represent the learned model within a prescribed distortion in the generalization gap, yielding informative guarantees for deterministic algorithms over continuous hypothesis spaces. For margin-based prediction, we derive explicit bounds for linear and self-attention next-token predictors via noisy low-dimensional compression, revealing the roles of context length, model complexity, sample size, margin, and temporal mixing. Experiments on TinyStories and ETTh2 show that longer contexts can reduce both training and test losses, but typically reduce training loss more, enlarging the generalization gap. A complementary ETTh2 analysis identifies an effective predictive-memory scale near 24 hours, with no statistically supported improvement beyond this scale, offering a plausible explanation for test-performance saturation at larger contexts.