User Model Extraction via Belief Self-Distillation

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of detecting and causally intervening upon implicit user beliefs encoded within large language models. We propose a belief self-distillation framework that learns compact, readable, and writable user representations from a frozen teacher model in an unsupervised manner. By unifying linear probing with causal tracing, this work reveals shared representational geometries across models and elucidates the causal influence of user intent on model refusal behavior. Experimental results demonstrate that the proposed method faithfully recovers user beliefs, while hidden-state-guided interventions significantly outperform conventional baselines. These findings offer new perspectives for AI safety and alignment research by enabling more precise identification and manipulation of latent user beliefs governing model behavior.
📝 Abstract
Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The frozen LLM acts as its own teacher, distilling beliefs from natural conversations without external annotations. Unlike conventional probing, BSD isolates not only information present in activations, but a state whose causal role can be directly tested. Across multiple model families, BSD faithfully recovers user beliefs and enables substantially stronger interventions than matched hidden-state steering. Crucially, we find that refusal depends not only on the request, but on the model's inferred user intent: changing this belief alters refusal while holding the request fixed. We further uncover a striking cross-model regularity: independently trained LLMs converge on a shared geometry for representing their users. Together, these results reveal implicit user models as readable and causally writable internal states with direct implications for AI safety, shaping how models condition safety decisions on whom they believe they are interacting with.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
User Model Extraction
Causal Probing
AI Safety
Implicit Beliefs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Belief Self-Distillation
Causal Probing
User Model Extraction
Hidden-State Steering
Cross-Model Regularity
🔎 Similar Papers
No similar papers found.