🤖 AI Summary
This study addresses the challenge in autoregressive pretraining where tokens simultaneously serve as prediction targets and contextual inputs, making it difficult to disentangle these dual roles and accurately assess underlying learning mechanisms. Adopting a "role disentanglement" perspective, this work isolates the two functions of tokens through controlled data corruption experiments. The findings reveal a reversal phenomenon between prediction difficulty and context degradation: while noise reduces prediction loss, it exacerbates contextual degradation, exposing a critical limitation in existing generative paradigms where the contextual role remains independently unverified. By elucidating the distinct contributions of each token function, this research provides a new theoretical foundation for precisely understanding and controlling what language models learn from individual tokens.
📝 Abstract
Autoregressive pretraining increasingly draws on heterogeneous data, making it important to understand how a model learns from an individual token. The next-token prediction objective naturally identifies a token's contribution with its own loss. However, each token is not only a prediction target but also context for what follows. Using controlled corruption, we decouple these two roles and find a reversal: making a noisy token easier to predict reduces its damage as a target but increases it as context. The same decoupling helps explain text generated by language models: generation selects each token by its fit to the prefix, while its role as context is never tested against an independently determined continuation, because that continuation is generated to fit it. At known corrupted positions, acting through the context can reduce damage that removing the token's own loss does not. Understanding and controlling what a model learns from a token therefore requires decoupling its roles.