🤖 AI Summary
This study investigates the distinction between the causal use and mere decodability of internal representations in Transformers when performing hierarchical tasks. By integrating representation probing, attention masking, and residual stream subspace ablation on Dyck languages and templated natural language tasks, the work provides the first clear separation between the decodability of hierarchical information and its causal role in model computation. The findings reveal that although hierarchical signals are widely decodable across representations, the model relies selectively on specific mechanisms—such as attention to stack-top positions—to handle long-range dependencies. Moreover, ablating low-dimensional residual subspaces has negligible impact on performance, demonstrating that decodability does not imply causal utilization. This work thus uncovers a critical gap between the readability of internal representations and the actual mechanisms driving model reasoning.
📝 Abstract
When trained on tasks requiring an understanding of hierarchical structure, transformers have been found to represent this hierarchy in distinct ways: in the geometry of the residual stream, and in stack-like attention patterns maintaining a last-in, first-out ordering. However, it remains unclear whether these representations are causally used or merely decodable. We examine this gap in transformers trained on the Dyck language (a formal language of balanced bracket sequences), where the hierarchical ground truth is explicit. By probing and intervening on the residual stream and attention patterns, we find that depth, distance, and top-of-stack signals are all decodable, yet their causal roles diverge. Specifically, masking attention to the true top-of-stack position causes a sharp drop in long-distance accuracy, while ablating low-dimensional residual stream subspaces has comparatively little effect. These results, which extend to a templated natural language setting, suggest that even in a controlled setting where the relevant hierarchical variables are known, decodability alone does not imply causal use.