"Many Are My Names": The Anatomy of the Assistant and Its Personas via Sparse Autoencoders

📅 2026-08-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how large language models internally represent distinct speaker identities—such as the default assistant persona, role-played personas, and narrative characters—and examines their similarities and differences. By constructing a user–model dialogue dataset and applying sparse autoencoders to extract representations at turn boundaries and pronoun positions, the authors design a multi-stage filtering framework to analyze representational structures under controlled generation settings. The findings reveal that the assistant and role-played personas share a common core representation that progressively diverges across layers, whereas narrative characters lack this shared core. The work introduces an “immersive simulation mode” that effectively discriminates between generation types, successfully identifies the assistant’s core features and their evolution during role-playing, and demonstrates that the model may spontaneously enter an immersive state even in default generation.
📝 Abstract
How a language model internally represents who is speaking, the Assistant, an assigned roleplay persona, or a narrated story character, remains underexplored. We study speaker representations using a dataset of user-expressed emotional text and corresponding model responses. We decompose three generation settings (Assistant, Roleplay, and Story) into sparse autoencoder features extracted at turn-boundary and pronoun-token positions and selected through a filtering pipeline for different depths. We characterize each surviving feature through its steering effects and activation distribution. Our main finding is that the Assistant and roleplay personas are not independent alternatives: personas retain the Assistant-associated feature core while progressively differentiating from it across layers, starting from operational machinery towards behavioral and stylistic features. Meanwhile, generated story characters lack the Assistant-associated core. Both Story and Roleplay can be distinguished from the Assistant with Immersive Simulation Mode. However, the Assistant can sometimes enter or slowly drift into it even in the default setting.
Problem

Research questions and friction points this paper is trying to address.

speaker representation
language model
personas
roleplay
story characters
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Autoencoders
Speaker Representation
Roleplay Persona
Immersive Simulation Mode
Feature Decomposition
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Adelaide Danilov
Department of Computer Science, Faculty of Science, Technology and Medicine, University of Luxembourg
A
Aria Nourbakhsh
Department of Computer Science, Faculty of Science, Technology and Medicine, University of Luxembourg
O
Oleksandr Marchenko Breneur
Department of Computer Science, Faculty of Science, Technology and Medicine, University of Luxembourg
Salima Lamsiyah
Salima Lamsiyah
NLP-Machine Learning Researcher, Luxembourg University
NLPMachine LearningDeep LearningTransfer LearningLLM