🤖 AI Summary
This study addresses the prohibitive computational cost of exhaustively interpreting natural language autoencoders by investigating token position prioritization for large language model (LLM) safety auditing. By comparing model-derived computational signals with chat-structural features, we demonstrate that a ranking algorithm relying solely on conversational structure achieves superior position selection compared to computationally intensive alternatives. Furthermore, pre-trained language models are leveraged to effectively recover lexical information obscured during fine-tuning. Experimental results indicate that this approach maintains near-complete threat detection rates across most datasets while interpreting merely 5% of token positions, thereby substantially reducing the overhead associated with LLM safety auditing.
📝 Abstract
Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across $4.7$ million explanations on prompt injection and concealment. We compare signals from model computation with a ranker trained only on chat structure. Chat structure usually selects more relevant explanations than the computational signals, without requiring a model forward pass for position selection. On three of four datasets, explaining just $5\%$ of positions retains nearly all of the success rate from explaining every position, where success means obtaining an explanation about the threat. The benefit varies with the audit task. We also show that pretrained verbalizers recover words that models have learned to conceal through fine-tuning, without additional verbalizer training. These results identify where auditors can concentrate explanation generation and show that useful explanations can extend beyond the model a verbalizer was trained to describe.