The Attribution Blind Spot: Layerwise Trajectory Diagnostics for Source Reliance in Retrieval-Augmented Language Models

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue wherein retrieval-augmented models match relevant documents yet fail to rely on source information, investigating internal state properties that define effective intervention directions. Methodologically, it leverages controlled knowledge conflicts to render source selection observable, employing latent trajectory shift (LTS) and principal component analysis to disentangle training exposure from behavioral choices. The findings reveal a decoupling between diagnostic and control representations: the magnitude of state change serves as a stronger predictor, whereas the equal-norm signed PC1 direction functions as a superior controller capable of precisely modulating source preferences while preserving non-target behaviors. Furthermore, these frozen directions demonstrate transferability across datasets.
📝 Abstract
A retrieval-augmented model can match a document without relying on it. Controlled knowledge conflicts make source choice observable and let us ask a second question that prediction alone cannot answer: which internal-state properties define useful intervention directions? We study paired hidden-state changes with Latent Trajectory Shift (LTS), a signed projection onto a training-fitted first principal component (PC1), and keep verified training exposure separate from behavioral source choice. Across the evaluated conflicts, state-change magnitude is often the stronger predictor, whereas signed PC1 is the stronger selective controller: equal-norm interventions change source preference while better preserving non-target behavior, and the frozen direction transfers across the tested datasets and aligned model pairs. A same-system OLMo study further combines positive choice and control results with inconclusive exposure detection at the achieved power. The central result is a separation: representations that diagnose what a model will choose need not be the representations that best control that choice.
Problem

Research questions and friction points this paper is trying to address.

Retrieval-Augmented Language Models
Source Reliance
Knowledge Conflicts
Internal Representations
Latent Trajectory Shift
Innovation

Methods, ideas, or system contributions that make the work stand out.

Retrieval-Augmented Language Models
Latent Trajectory Shift
Principal Component Analysis
Knowledge Conflicts
Mechanistic Interpretability
🔎 Similar Papers
No similar papers found.