Dissociating Decodability and Causal Use in Bracket-Sequence Transformers

📅 2026-04-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the distinction between the causal use and mere decodability of internal representations in Transformers when performing hierarchical tasks. By integrating representation probing, attention masking, and residual stream subspace ablation on Dyck languages and templated natural language tasks, the work provides the first clear separation between the decodability of hierarchical information and its causal role in model computation. The findings reveal that although hierarchical signals are widely decodable across representations, the model relies selectively on specific mechanisms—such as attention to stack-top positions—to handle long-range dependencies. Moreover, ablating low-dimensional residual subspaces has negligible impact on performance, demonstrating that decodability does not imply causal utilization. This work thus uncovers a critical gap between the readability of internal representations and the actual mechanisms driving model reasoning.

Technology Category

Knowledge Representation and Reasoning: Action, Change, and CausalityReasoning under Uncertainty: CausalityMachine Learning: Representation Learning

Application Category

Search and Retrieval-Augmented AI: Web query analysis, representation and understandingSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsGraph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphs
📝 Abstract
When trained on tasks requiring an understanding of hierarchical structure, transformers have been found to represent this hierarchy in distinct ways: in the geometry of the residual stream, and in stack-like attention patterns maintaining a last-in, first-out ordering. However, it remains unclear whether these representations are causally used or merely decodable. We examine this gap in transformers trained on the Dyck language (a formal language of balanced bracket sequences), where the hierarchical ground truth is explicit. By probing and intervening on the residual stream and attention patterns, we find that depth, distance, and top-of-stack signals are all decodable, yet their causal roles diverge. Specifically, masking attention to the true top-of-stack position causes a sharp drop in long-distance accuracy, while ablating low-dimensional residual stream subspaces has comparatively little effect. These results, which extend to a templated natural language setting, suggest that even in a controlled setting where the relevant hierarchical variables are known, decodability alone does not imply causal use.
Problem

Research questions and friction points this paper is trying to address.

decodability
causal use
hierarchical structure
transformers
Dyck language
Innovation

Methods, ideas, or system contributions that make the work stand out.

decodability
causal use
hierarchical structure
attention intervention
residual stream
🔎 Similar Papers
2024-10-02arXiv.orgCitations: 2
A
Aryan Sharma
Yale University
C
Cutter Dawes
Independent
S
Shivam Raval
Harvard University