LUMOS: Tracing Parametric Knowledge from Training Data to Behavioral Outputs in LLMs

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of distinguishing generalization from memorization in existing LLM knowledge analyses due to the lack of training data verification. We propose LUMOS, a framework that leverages the fully transparent corpus of OLMo 2 to incorporate training data exposure as an evaluation axis for the first time. By integrating causal tracing with chain-of-thought prompting, LUMOS constructs a complete causal chain from data exposure to behavioral output, enabling falsifiable knowledge diagnostics. Our findings reveal a significant dissociation in rare facts, which exhibit high internal encoding yet low behavioral expression, and demonstrate that self-reflection mechanisms fail on unseen data. This work elucidates the discrepancy between internal representations and external expressions, providing a rigorous new paradigm for knowledge attribution in large language models.
📝 Abstract
Current analyses of LLMs' parametric knowledge are largely output-centric, drawing conclusions about what a model knows without verifying what it was actually trained on. This leaves fundamental questions, such as whether a correct response reflects genuine generalization or rote memorization, grounded in speculation rather than evidence. To resolve these ambiguities, we introduce LUMOS, a diagnostic framework that traces knowledge along the causal chain from training-data exposure to behavioral output, leveraging OLMo 2 with its fully transparent training corpus. By grounding analysis in verified exposure, we reveal that models internally encode rare facts with high separability (84%) yet fail to express them behaviorally (54%), though this retrieval gap narrows with scale. Furthermore, when models are asked to self-reflect on their own answers, they perform reliably on trained content (83%) but drop to random-baseline levels (49%) on unseen content. This collapse persists even under chain-of-thought prompting, which inflates confidence signals rather than improving calibration. Collectively, these findings demonstrate that incorporating the training-data axis into LLM evaluation transforms speculative diagnoses into verifiable claims, and we advocate that this axis should be a standard component of knowledge assessment in LLMs.
Problem

Research questions and friction points this paper is trying to address.

parametric knowledge
large language models
training data exposure
knowledge tracing
model evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Parametric Knowledge Tracing
Causal Chain Analysis
Diagnostic Framework
Training Data Exposure
Self-Reflection Calibration
🔎 Similar Papers