🤖 AI Summary
This work investigates large language models’ (LLMs) sensitivity to temporal structure in in-context learning, specifically examining event retrieval biases under temporal–semantic decoupling. To isolate temporal effects, we design controlled prompt sequences: token positions of repeated elements are fixed, while all other tokens are randomly permuted—thereby eliminating semantic confounds. Using this setup, we systematically evaluate temporal preferences in both Transformer and state-space models (SSMs). Results show that both architectures exhibit strong recency and primacy effects—preferentially recalling tokens near sequence boundaries—and consistently prioritize prediction immediately following repeated tokens, while recall reliability drops markedly at middle positions. Ablation studies confirm that this behavior stems from induction heads rather than architecture-specific properties. To our knowledge, this is the first study to reveal a shared temporal inductive bias across LLMs under semantically isolated conditions, providing novel empirical evidence and an interpretable mechanistic account of time-related biases in in-context learning.
📝 Abstract
In-context learning is governed by both temporal and semantic relationships, shaping how Large Language Models (LLMs) retrieve contextual information. Analogous to human episodic memory, where the retrieval of specific events is enabled by separating events that happened at different times, this work probes the ability of various pretrained LLMs, including transformer and state-space models, to differentiate and retrieve temporally separated events. Specifically, we prompted models with sequences containing multiple presentations of the same token, which reappears at the sequence end. By fixing the positions of these repeated tokens and permuting all others, we removed semantic confounds and isolated temporal effects on next-token prediction. Across diverse sequences, models consistently placed the highest probabilities on tokens following a repeated token, but with a notable bias for those nearest the beginning or end of the input. An ablation experiment linked this phenomenon in transformers to induction heads. Extending the analysis to unique semantic contexts with partial overlap further demonstrated that memories embedded in the middle of a prompt are retrieved less reliably. Despite architectural differences, state-space and transformer models showed comparable temporal biases. Our findings deepen the understanding of temporal biases in in-context learning and offer an illustration of how these biases can enable temporal separation and episodic retrieval.