🤖 AI Summary
This study investigates whether weight sharing in recurrent Transformers forces identical computations across iterations and how hidden states route distinct operations. Using graph traversal as a testbed, the authors employ activation patching, linear probing, and intermediate supervision to conduct mechanistic interpretability analysis. The findings reveal that recurrent Transformers do not merely repeat fixed algorithms; instead, they modify input hidden states to guide shared layers toward executing different transformations. Specifically, a linear layer J functions as a causal pathway—mediated by attention—that governs transformation selection without altering shared parameters, effectively selecting from backbone capabilities rather than inventing new algorithms. This work elucidates the dynamic routing mechanism through which hidden states regulate shared computation in recurrent architectures and demonstrates that training supervision strategies fundamentally shape the set of available transformations.
📝 Abstract
Looped Transformers repeatedly apply the same set of Transformer layers, giving them a recurrent architecture for latent computation. Their strong performance on iterative reasoning and length-generalization tasks suggests an appealing explanation: recurrence may provide an inductive bias that lets the model reuse a learned algorithm across loops. However, weight sharing alone does not imply that every loop performs the same operation. This raises a basic question: is each loop actually repeating the same computation, and if not, what routes the shared parameters to different operations?
We study this question using graph walks as a test case. In the model's native trajectories, decoded predictions can advance by different numbers of graph steps or remain at a reached target, showing that recurrent progress need not follow a fixed one-loop-one-step pattern. We then show that a frozen loop can be steered toward different transitions by modifying its entering hidden state: a learned linear layer $J$ selects the desired transition without changing the shared Transformer layers.
To test how this steering works, we use activation patching and find that attention patterns can recover its effects and switch the selected transition. Across five matched pairs of graph models, changing intermediate supervision during backbone training changes which transitions $J$ can induce. This suggests that $J$ selects computations learned by the backbone rather than creating new algorithms. Together, these results show that the hidden state can control shared computation, with attention routing as a causal pathway.