Beyond Semantics: How Temporal Biases Shape Retrieval in Transformer and State-Space Models

📅 2025-10-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates large language models’ (LLMs) sensitivity to temporal structure in in-context learning, specifically examining event retrieval biases under temporal–semantic decoupling. To isolate temporal effects, we design controlled prompt sequences: token positions of repeated elements are fixed, while all other tokens are randomly permuted—thereby eliminating semantic confounds. Using this setup, we systematically evaluate temporal preferences in both Transformer and state-space models (SSMs). Results show that both architectures exhibit strong recency and primacy effects—preferentially recalling tokens near sequence boundaries—and consistently prioritize prediction immediately following repeated tokens, while recall reliability drops markedly at middle positions. Ablation studies confirm that this behavior stems from induction heads rather than architecture-specific properties. To our knowledge, this is the first study to reveal a shared temporal inductive bias across LLMs under semantically isolated conditions, providing novel empirical evidence and an interpretable mechanistic account of time-related biases in in-context learning.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsCognitive Modeling & Cognitive Systems: Conceptual Inference and Reasoning

Application Category

Search and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
In-context learning is governed by both temporal and semantic relationships, shaping how Large Language Models (LLMs) retrieve contextual information. Analogous to human episodic memory, where the retrieval of specific events is enabled by separating events that happened at different times, this work probes the ability of various pretrained LLMs, including transformer and state-space models, to differentiate and retrieve temporally separated events. Specifically, we prompted models with sequences containing multiple presentations of the same token, which reappears at the sequence end. By fixing the positions of these repeated tokens and permuting all others, we removed semantic confounds and isolated temporal effects on next-token prediction. Across diverse sequences, models consistently placed the highest probabilities on tokens following a repeated token, but with a notable bias for those nearest the beginning or end of the input. An ablation experiment linked this phenomenon in transformers to induction heads. Extending the analysis to unique semantic contexts with partial overlap further demonstrated that memories embedded in the middle of a prompt are retrieved less reliably. Despite architectural differences, state-space and transformer models showed comparable temporal biases. Our findings deepen the understanding of temporal biases in in-context learning and offer an illustration of how these biases can enable temporal separation and episodic retrieval.
Problem

Research questions and friction points this paper is trying to address.

Investigating temporal biases in transformer and state-space models
Isolating temporal effects on next-token prediction in LLMs
Analyzing episodic retrieval capabilities across different model architectures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Models isolate temporal effects via fixed token positions
Ablation links transformer bias to induction heads
State-space and transformer models show comparable biases
Indiana University Bloomington
A
Anooshka Bajaj
Department of Computer Science, Indiana University Bloomington
D
Deven Mahesh Mistry
Department of Computer Science, Indiana University Bloomington
Sahaj Singh Maini
Sahaj Singh Maini
Indiana University
Machine LearningCognitive ScienceComputational Neuroscience
Y
Yash Aggarwal
Department of Computer Science, Indiana University Bloomington
Zoran Tiganj
Zoran Tiganj
Department of Computer Science, Department of Psychological and Brain Sciences, Indiana University
Artificial IntelligenceCognitive ScienceComputational Neuroscience