Emergent Stack Representations in Modeling Counter Languages Using Transformers

📅 2025-02-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
It remains unclear whether Transformers inherently develop stack-like representations when modeling counting languages—despite lacking explicit stack mechanisms. Method: We construct four synthetic counting languages, each equivalent to varying stack depth, and train standard autoregressive Transformers on them. Using regression probes applied to hidden states, we quantify how well stack depth is encoded at each token position. Contribution/Results: We provide the first empirical evidence that Transformers—trained solely for next-token prediction—encode theoretical stack depth with high fidelity in their internal representations (mean R² > 0.92). This demonstrates that stack-structured representations can emerge spontaneously without architectural constraints or auxiliary objectives. Our findings reveal an intrinsic stack-shaped inductive bias in Transformers, offering a new paradigm for mechanistic interpretability and grounding circuit-level analysis in robust, probe-based evidence.

Technology Category

Machine Learning: Representation LearningNatural Language Processing: (Large) Language ModelsKnowledge Representation and Reasoning: Computational Complexity of Reasoning

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingWeb Mining and Content Analysis: Large pretrained models with web dataSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Transformer architectures are the backbone of most modern language models, but understanding the inner workings of these models still largely remains an open problem. One way that research in the past has tackled this problem is by isolating the learning capabilities of these architectures by training them over well-understood classes of formal languages. We extend this literature by analyzing models trained over counter languages, which can be modeled using counter variables. We train transformer models on 4 counter languages, and equivalently formulate these languages using stacks, whose depths can be understood as the counter values. We then probe their internal representations for stack depths at each input token to show that these models when trained as next token predictors learn stack-like representations. This brings us closer to understanding the algorithmic details of how transformers learn languages and helps in circuit discovery.
Problem

Research questions and friction points this paper is trying to address.

Transformer Models
Counter Languages
Internal Mechanisms
Innovation

Methods, ideas, or system contributions that make the work stand out.

Transformer Models
Counter Languages
Stack-like Processing
🔎 Similar Papers
No similar papers found.