Selecting The Most Informative Tokens in Natural Language Autoencoders

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive computational cost of exhaustively interpreting natural language autoencoders by investigating token position prioritization for large language model (LLM) safety auditing. By comparing model-derived computational signals with chat-structural features, we demonstrate that a ranking algorithm relying solely on conversational structure achieves superior position selection compared to computationally intensive alternatives. Furthermore, pre-trained language models are leveraged to effectively recover lexical information obscured during fine-tuning. Experimental results indicate that this approach maintains near-complete threat detection rates across most datasets while interpreting merely 5% of token positions, thereby substantially reducing the overhead associated with LLM safety auditing.
📝 Abstract
Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across $4.7$ million explanations on prompt injection and concealment. We compare signals from model computation with a ranker trained only on chat structure. Chat structure usually selects more relevant explanations than the computational signals, without requiring a model forward pass for position selection. On three of four datasets, explaining just $5\%$ of positions retains nearly all of the success rate from explaining every position, where success means obtaining an explanation about the threat. The benefit varies with the audit task. We also show that pretrained verbalizers recover words that models have learned to conceal through fine-tuning, without additional verbalizer training. These results identify where auditors can concentrate explanation generation and show that useful explanations can extend beyond the model a verbalizer was trained to describe.
Problem

Research questions and friction points this paper is trying to address.

natural language autoencoders
informative token selection
prompt injection
concealment
model auditing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Token Selection
Natural Language Autoencoders
Prompt Injection
Mechanistic Interpretability
Verbalizers
🔎 Similar Papers
No similar papers found.
F
Federico Torrielli
Department of Computer Science, University of Turin, Italy
G
Gianluca Barmina
Department of Mathematics and Computer Science, University of Southern Denmark, Denmark
A
Andrea Blasi Núñez
Department of Mathematics and Computer Science, University of Southern Denmark, Denmark
Amon Rapp
Amon Rapp
Associate Professor, Università degli Studi di Torino, Torino
Human-Computer InteractionHCIComputer-Supported Cooperative WorkCSCWHuman-Centered AI
Luigi Di Caro
Luigi Di Caro
Associate Professor
data miningnatural language processing
Peter Schneider-Kamp
Peter Schneider-Kamp
Professor of Computer Science, University of Southern Denmark
Artificial IntelligenceAutomated ReasoningDeclarative ProgrammingProgramming LanguagesSoftware Verification
L
Lukas Galke Poech
Department of Mathematics and Computer Science, University of Southern Denmark, Denmark