Identifying Temporal Features within Transcoders for Time Sensitive Factual Recall

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the misalignment in temporal-sensitive factual recall within large language models caused by contradictory training data. By leveraging transcoder circuit tracing to isolate MLP features, we conduct interpretability analyses on architectures including Gemma and LLaMA. This work constructs the first feature-level temporal recall graph, identifying three categories of critical nodes. Our findings reveal that temporal representations do not follow a simple linear pipeline but instead operate through a cross-layer, parallel, hybrid interaction mechanism, while also uncovering high-level temporal components. Ultimately, this project establishes the transcoder as a more comprehensive lens for temporal interpretability, providing a theoretical foundation for targeted interventions.
📝 Abstract
Large Language Models (LLMs) suffer from temporal misalignment, often due to the contradictory nature of their training corpora. While current mitigation strategies rely on computationally expensive fine-tuning or context-heavy retrieval augmented generation (RAG), the internal mechanisms governing time-sensitive recall remain under-explored. Unlike prior studies that identify temporal components such as attention heads and MLP layers, we provide the first feature-level map of temporal recall by isolating individual MLP features via transcoder circuit tracing. We identify three node categories (common temporal, common to the year, and chrono-semantic) which interact to generate a temporal filter during factual recall. By analysing Gemma 2 2B, LLaMA 3.2 1B, and Qwen3-4B, we show that these features do not follow a simple linear pipeline but represent time through a parallel and mixed syntactic-semantic interplay across layers. We additionally discover a class of higher-layer temporal components invisible to existing EAP-IG methods, establishing transcoders as a more complete lens for temporal interpretability in time-sensitive factual recall. These findings present MLP components for potential targeted interventions in time-sensitive factual recall
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Temporal Misalignment
Time-sensitive Factual Recall
Mechanistic Interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Transcoder Circuit Tracing
Temporal Interpretability
Feature-level Analysis
Time-sensitive Factual Recall
MLP Features
S
Sanjay Govindan
University of New South Wales, Sydney
Y
Yang Song
University of New South Wales, Sydney
Maurice Pagnucco
Maurice Pagnucco
Professor, The University of New South Wales
Artificial Intelligence