🤖 AI Summary
This study addresses the problem of directed feature learning for predefined concepts in language models, bridging a gap in sparse autoencoders for hypothesis-driven scenarios. By systematically evaluating combinations of signals—such as activations and gradients—with various estimators, this work proposes four novel methods for directed feature extraction and benchmarks them against contrastive mean, 1D autoencoders, and sparse autoencoders. Results reveal that concept detection and causal intervention are predominantly governed by activations and gradients, respectively, exhibiting complementary properties. Furthermore, the quality of directed features depends critically on the joint selection of signal and estimator. By delineating the optimal application scenarios for each approach, this research provides systematic guidance for controllable feature extraction in language models.
📝 Abstract
Model-internal features can be studied through both their ability to identify a specified concept and their causal effect when manipulated, e.g., through steering or weight editing. A prominent approach to feature learning is Sparse Autoencoders (SAEs), which learn broad feature dictionaries whose relation to particular concepts is typically identified post hoc. However, many interpretability questions are instead hypothesis-driven and concern a concept specified in advance. We study this setting as targeted feature learning, where a single feature is constructed for such a predefined concept. We present a controlled comparison across three model signals (activation values, activation gradients, and parameter gradients) and two estimators (contrastive mean and a learned one-dimensional encoder-decoder), yielding six targeted methods, with CAA and GRADIEND as existing instances and four new methods covering the remaining combinations. We compare these methods against pretrained SAEs across 15 tasks and three language models, evaluating both detection and causal intervention. Across models, the strongest detection performance is achieved by contrastive activation value methods, whereas the strongest intervention performance is achieved by gradient-based methods. Overall, our results show that targeted feature quality depends jointly on the model signal and estimator, with detection and intervention capturing complementary properties.