π€ AI Summary
This study reveals a βconcept-specific blind spotβ in fine-tuned Activation Oracles (AOs): even when a target concept is consistently present in their training data, AOs may systematically fail to attend to it, effectively acting as βanti-readers.β The authors train AOs on a taboo-word guessing task to interpret internal activations of another model and combine this with LogitLens analysis, layer ablation, and representational decodability assessments. They find that while the target concept remains decodable within the AOβs internal representations, the readout pathway fails to express it. This dissociation among behavioral leakage, representational decodability, and AO expressivity challenges the assumption that AOs serve as reliable interpretability interfaces and highlights potential reliability risks inherent in learned explanation tools.
π Abstract
Activation Oracles (AOs) are language models trained to answer natural-language questions about another model's internal activations. They offer a flexible interface for reading hidden information from model states, especially when relevant information is internally represented but absent or incomplete in visible behavior. However, AOs are themselves learned systems: their answers are shaped by training data, objectives, and learned reporting behavior, rather than being neutral readouts of represented information. We study this in a controlled Taboo Word Guessing setting, where subject models are fine-tuned to internally use a hidden concept while avoiding direct disclosure. Contrary to the expectation that an AO trained on such a subject becomes a specialist reader, we find that fine-tuned AOs can become concept-specific anti-readers: they selectively fail to recover the concept persistently present during their own training. This failure is not simply explained by absence of the concept from the subject or oracle representations: the target remains decodable inside the oracle, while LogitLens and layer-ablation analyses indicate that the failure arises in the AO readout pathway. Our results show that behavioral leakage, representation-level decodability, and AO-verbalizability can come apart, raising a reliability concern for learned interpretability interfaces.