🤖 AI Summary
This study investigates how recurrent Transformers achieve chain-of-thought generalization beyond their training lengths, examining the underlying computational principles and failure boundaries. By comparing Memory-Recurrent (MR) and Diagonal-Recurrent (DR) architectures, the authors employ mechanistic interpretability techniques, including attention analysis, intermediate-state decoding, and causal interventions. The findings reveal that "live consumption" and "fresh aging" residual stream directions govern information availability, demonstrating that models can leverage task structure for generalization without performing faithful step-by-step reasoning. Furthermore, this work elucidates the differences in state propagation between the two architectures and identifies the root causes of depth degradation, ultimately establishing universal representational principles that disentangle encoded content from computational states.
📝 Abstract
Looped Transformers can generalize to reasoning chains longer than those encountered during training, but the computations enabling this behavior and limiting its extent remain unclear. We mechanistically compare two looped-Transformer configurations, which we call the Matched-Recurrence Looped Transformer (MR-Loop) and Decoupled-Recurrence Looped Transformer (DR-Loop), reflecting their respective recurrence-training schemes. We evaluate polynomial iteration, finite-state composition, and knowledge-graph traversal using detailed mechanistic analysis. Attention analysis, intermediate-state decoding, and causal interventions reveal distinct mechanisms learned under final-answer supervision. MR-Loop updates an intermediate state at a fixed readout while advancing relation selection through adjacent-token interactions and a transferable progress cue. DR-Loop instead propagates intermediate states across relation positions, forming an advancing computational frontier. However, both mechanisms become unreliable at greater depths: MR-Loop exhibits degradation of its readout state and progress cues, while DR-Loop exhibits declining reliability of state propagation. Limited self-correction allows local errors to persist and compound. Across both models, we uncover a common representational principle: recurrent states encode not only task-relevant content but also its computational status, whether that content remains in a form that can support subsequent computation. Transferable live-consumed and fresh-aged residual directions causally control whether represented information can participate in subsequent computation, including beyond the training horizon. We further show that length generalization need not rely on faithful step-by-step reasoning, as Looped Transformers can exploit task structure without explicitly representing every intermediate state.