Tracking States or Tracking Cosets? An Algebraic Account of Learned State Tracking

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the algebraic mechanisms underlying neural network state-tracking accuracy, focusing on whether models learn full group states or quotient structures. Through a group element product prediction task, it integrates group theory, algebraic topology, and internal representation geometry to contrast the encoding strategies of Transformers and recurrent neural networks (RNNs). The analysis demonstrates that optimal order-blind accuracy converges to the reciprocal of the abelianization class size. Furthermore, it reveals that RNNs effectively encode right cosets of non-normal subgroups, whereas Transformers cannot, and identifies approximate dodecahedral geometric structures within low-dimensional subspaces. By elucidating the intrinsic relationship between partial accuracy and learning stages, this work provides novel perspectives for understanding algebraic generalization in neural models.
📝 Abstract
State tracking requires composing a sequence of updates, but accuracy alone does not reveal what a model has learned. We study neural networks trained to predict the running product of group elements. We identify quotient solutions in Transformers, where models recover the quotient class while predicting nearly uniformly among its members. The reciprocal of class size predicts partial accuracy without a fitted parameter, extending parity-based accounts to non-parity quotients. Our baseline Transformers' predictions change little under prefix reordering beyond the exact-tracking frontier. We prove that, for finite groups under uniform i.i.d. full-group inputs, optimal order-blind exact accuracy converges to the reciprocal of abelianization class size as prefix length grows, consistent with the observed abelianization plateaus. Sequential updates permit more: any partition into right cosets of a subgroup, normal or not, survives sequential updates. In our census of standard Transformers, every recovered coset partition comes from a normal subgroup, whereas parameter-matched recurrent networks pass through both normal and non-normal right-coset stages during training. On $A_5$, we identify low-dimensional subspaces of the recurrent state that encode non-normal cosets. In the three-dimensional cases, coset mean vectors form approximate dodecahedra, and swapping the state components in these subspaces transfers the donor's coset state through a shared input suffix. Our results connect partial accuracy, learning stages, and internal computation through the subgroup cosets that models learn to track.
Problem

Research questions and friction points this paper is trying to address.

state tracking
group theory
coset partitions
neural network interpretability
Transformers
Innovation

Methods, ideas, or system contributions that make the work stand out.

state tracking
group theory
coset partition
transformer
recurrent networks