Transformer Heads Looking for Order
研究解决了使用Transformer判断比特序列是否有序的问题,证明单头单层Transformer无法完成,而双头单层可以实现,模型包含输出MLP。
研究解决了使用Transformer判断比特序列是否有序的问题,证明单头单层Transformer无法完成,而双头单层可以实现,模型包含输出MLP。
This work establishes fundamental limitations of non-negative kernel attention mechanisms, showing that their feature dimension must grow exponentially even to handle contexts as short as three tokens. Focusing on the exact modeling of Boolean inputs under the Min-IP task, the authors employ rank analysis, information-theoretic lower bounds, and constructive counterexamples within a framework of causal queries and position-dependent mappings. Their key contribution is the first proof that standard Softmax attention solves this task with only linear feature dimensionality, whereas non-negative kernel attention heads require at least $2^{\Omega(m)}$ dimensions. Furthermore, they derive an information transmission lower bound for multi-head, multi-layer models operating over finite alphabets.
This work investigates the information bottleneck inherent in modern deep learning architectures when performing indexing tasks, with a particular focus on performance disparities between indexing at the beginning versus the end of a sequence. By introducing the notion of causal complexity, the authors theoretically demonstrate that low-parameter RNNs, state space models (SSMs), and masked linear-attention Transformers cannot solve the tail-indexing task within a constant number of layers, whereas a Softmax-based Transformer can accomplish it in a single layer. Combining theoretical analysis under infinite precision with empirical validation for sequence lengths up to \( n \leq 64 \), the study shows that architectures with low causal complexity learn indexing tasks efficiently, while those with higher causal complexity degrade significantly as sequence length increases.
This study addresses the problem of achieving error-bounded language generation in polynomial time, with a focus on specific classes of Boolean functions. The work proposes a polynomial-time framework grounded in combinatorial game theory and establishes, for the first time, that all monotone Boolean functions with a polynomial number of maximal terms—encompassing every monotone Boolean function computable by a decision tree of polynomial size—admit error-bounded language generation within this framework. Furthermore, the paper demonstrates the polynomial-time learnability under bounded error of variable parity functions, conjunctions of literals, and a broad class of monotone Boolean functions, thereby significantly extending the known boundaries for both generation and learning of Boolean functions under constrained error conditions.
This study investigates how Transformers leverage attention mechanisms to generalize on structured multi-hop reasoning tasks, with a particular focus on their divergent performance in long-sequence extrapolation. By training GPT-J models on equivalent numeric and alphabetic tasks and employing controlled experiments, attention head behavior classification, and geometric analysis of Rotary Position Embeddings (RoPE), the work provides the first clear distinction and theoretical characterization of positional versus symbolic attention heads in terms of their computational roles. The authors introduce a “discrepancy” metric that quantitatively demonstrates the superior extrapolation robustness of symbolic mechanisms over positional ones. Furthermore, they establish that the presence of purely typed attention heads is critical for successful learning—a finding consistently validated across both controlled setups and real-world models.
研究解决了使用Transformer判断比特序列是否有序的问题,证明单头单层Transformer无法完成,而双头单层可以实现,模型包含输出MLP。
This work establishes fundamental limitations of non-negative kernel attention mechanisms, showing that their feature dimension must grow exponentially even to handle contexts as short as three tokens. Focusing on the exact modeling of Boolean inputs under the Min-IP task, the authors employ rank analysis, information-theoretic lower bounds, and constructive counterexamples within a framework of causal queries and position-dependent mappings. Their key contribution is the first proof that standard Softmax attention solves this task with only linear feature dimensionality, whereas non-negative kernel attention heads require at least $2^{\Omega(m)}$ dimensions. Furthermore, they derive an information transmission lower bound for multi-head, multi-layer models operating over finite alphabets.
This work investigates the information bottleneck inherent in modern deep learning architectures when performing indexing tasks, with a particular focus on performance disparities between indexing at the beginning versus the end of a sequence. By introducing the notion of causal complexity, the authors theoretically demonstrate that low-parameter RNNs, state space models (SSMs), and masked linear-attention Transformers cannot solve the tail-indexing task within a constant number of layers, whereas a Softmax-based Transformer can accomplish it in a single layer. Combining theoretical analysis under infinite precision with empirical validation for sequence lengths up to \( n \leq 64 \), the study shows that architectures with low causal complexity learn indexing tasks efficiently, while those with higher causal complexity degrade significantly as sequence length increases.
This study addresses the problem of achieving error-bounded language generation in polynomial time, with a focus on specific classes of Boolean functions. The work proposes a polynomial-time framework grounded in combinatorial game theory and establishes, for the first time, that all monotone Boolean functions with a polynomial number of maximal terms—encompassing every monotone Boolean function computable by a decision tree of polynomial size—admit error-bounded language generation within this framework. Furthermore, the paper demonstrates the polynomial-time learnability under bounded error of variable parity functions, conjunctions of literals, and a broad class of monotone Boolean functions, thereby significantly extending the known boundaries for both generation and learning of Boolean functions under constrained error conditions.
This study investigates how Transformers leverage attention mechanisms to generalize on structured multi-hop reasoning tasks, with a particular focus on their divergent performance in long-sequence extrapolation. By training GPT-J models on equivalent numeric and alphabetic tasks and employing controlled experiments, attention head behavior classification, and geometric analysis of Rotary Position Embeddings (RoPE), the work provides the first clear distinction and theoretical characterization of positional versus symbolic attention heads in terms of their computational roles. The authors introduce a “discrepancy” metric that quantitatively demonstrates the superior extrapolation robustness of symbolic mechanisms over positional ones. Furthermore, they establish that the presence of purely typed attention heads is critical for successful learning—a finding consistently validated across both controlled setups and real-world models.