Score
Designs and implements transformer architectures that apply the same transformer block repeatedly (looped) across depth or time with tied weights, producing recurrent or weight‑shared models that reduce stored parameters. This work covers specifying the recurrence/unrolling scheme and state/attention carryover, adapting positional and normalization components, and choosing training strategies (e.g., backpropagation through recurrence or truncated unrolling) while analyzing the resulting trade‑offs in capacity, efficiency, and convergence.
Transformers exhibit inherent limitations in modeling long-range context, continual learning, and knowledge integration. To address these challenges, we propose a neuroscience-inspired memory-augmented unified framework that integrates multi-timescale memory, selective attention, and synaptic consolidation mechanisms—shifting from static caching to adaptive online learning. Methodologically, we design a synergistic architecture combining attention fusion, gating control, and associative retrieval, supported by a hybrid memory representation comprising parametric encoding, state-based internal representations, and explicit external memory. We further introduce a hierarchical buffering structure and a surprise-driven memory update strategy to mitigate capacity bottlenecks and catastrophic forgetting. Experiments demonstrate substantial improvements in long-sequence modeling stability and cross-task knowledge transfer. Our framework provides a scalable, biologically plausible pathway toward intelligent models capable of lifelong learning.
This work addresses the training instability and poor hyperparameter transferability in Looped Transformers, which arise from highly correlated residual updates across iterations due to weight sharing. The authors demonstrate that conventional depth-scaling strategies (e.g., ε = 1/√L) are insufficient in such recurrent architectures and reveal, for the first time, that residual correlations necessitate a stronger 1/N scaling with respect to the number of loop iterations N. They propose a composite scaling strategy, ε = λ/(N√L), which decouples the learning rate from N and makes it dependent solely on the effective depth L. Grounded in theoretical analysis and multi-layer loop block modeling, experiments confirm that this approach significantly enhances training stability, consistently outperforms 1/√N scaling across varying N, and enables direct transfer of large-N models without re-tuning hyperparameters.
This work addresses the limitations of conventional Transformers, which are constrained by finite effective depth, and recurrent models, which suffer from optimization instability and poor hardware efficiency despite their theoretically infinite depth. The authors propose a recurrent Transformer architecture that incorporates intra-layer recurrent attention, wherein each layer dynamically computes key-value pairs based on its own historical activations. This design significantly increases effective model depth while preserving the computational cost of standard autoregressive decoding. By unifying Transformer and token-level recurrence, the approach avoids training instability and achieves superior performance with fewer layers. Furthermore, a block-wise exact algorithm reduces HBM traffic from Θ(N²) to Θ(N log N), yielding an arithmetic intensity of Θ(N / log N). Experiments demonstrate that the model outperforms standard Transformer baselines of comparable size on C4 pretraining tasks with 150M and 300M parameters.
Increasing Transformer depth leads to exponential growth in parameter count, while existing recurrent methods perform coarse-grained, layer-level repetition without fine-grained control over computation. Method: We propose Intra-Layer Recurrence (ILR), a fine-grained recurrence mechanism that—within a single forward pass—selectively iterates core submodules (e.g., FFN or attention) multiple times inside a single Transformer layer, enabling dynamic state reuse without adding parameters, modifying architecture, or introducing auxiliary computational graphs. ILR employs a learnable iteration scheduling policy (e.g., allocating more iterations to earlier layers) to adaptively allocate compute resources. Contribution/Results: Evaluated on standard Transformer architectures for language modeling, ILR achieves comparable or superior performance to deeper baselines using significantly fewer parameters, thereby improving the trade-off between parameter efficiency and modeling capacity.
This paper addresses the limited recursive capability of Transformers in modeling long sequences. It systematically investigates two enhancement paradigms: depth-wise recursion (e.g., Universal Transformer) and chunked temporal recursion (e.g., Temporal Latent Bottleneck). The authors propose two key innovations: (1) a dynamic halting mechanism based on global mean activation, enabling adaptive computation steps; and (2) the first integration of depth-wise recursion into a chunked temporal framework, yielding a cross-paradigm fused architecture. Extensive evaluation on the Long-Range Arena (LRA) benchmark—including ListOps, Flip-Flop, and Logical Inference—demonstrates significant differences in the effectiveness of various recursive inductive biases. Results confirm that the dynamic halting mechanism jointly improves both modeling accuracy and computational efficiency. The implementation is publicly available.
Long-sequence modeling faces fundamental challenges including limited context length, difficulty in capturing long-range dependencies, and low efficiency in online learning. To address these, this work systematically reviews the resurgence of state-space models (SSMs) and recurrent computation, proposing a novel local forward-computation paradigm tailored for real-world online learning—thereby circumventing the temporal backtracking constraints inherent in standard backpropagation through time (BPTT). We introduce the first unified taxonomy encompassing both deep SSMs and large-context Transformers. Our framework integrates structured linear attention, enhanced RNN architectures, local recurrence mechanisms, and online optimization algorithms. The study rigorously clarifies the theoretical representational advantages and practical sequential reasoning benefits of recurrent modeling over alternatives. Collectively, this work delivers a scalable technical roadmap for low-latency, highly extensible long-sequence modeling.
This study investigates under what conditions weight-shared recurrent Transformers can learn and generalize genuine algorithmic behaviors, with a focus on the group word problem. Through controlled experiments combining SGD optimization, operator-priority curriculum learning, and causal ablation cone analysis, the work reveals that models tend to learn linear sequential algorithms, with their computational frontier governed by training budget according to \( v \sim n_{\text{train}} / T_{\text{train}} \) (R² = 0.99). The findings indicate that architectural priors dominate algorithm selection, and group order—not circuit complexity—constitutes the primary learning bottleneck. The paper proposes a convergence-rate-based stopping rule, a transferable mechanism, and introduces a novel head-measurement tool \( \tau(n,i) \). Full generalization is achieved on the A5 group, and results are successfully replicated on an easy-to-hard benchmark.
This work addresses the instability in recurrent Transformers caused by parameter sharing, particularly the mismatch between forward propagation and gradient aggregation when a module is accessed multiple times. To resolve this, the study introduces the access-aligned coefficient κ_R, which explicitly incorporates the number of parameter accesses into residual scaling design to model parameter update dynamics under recurrent depth. Through first-order perturbation analysis, the authors demonstrate that stable training requires increasing the scaling exponent in Post-LN DeepNorm from 1/4 to 1/2, leading to the proposed DeepLoop method. Experiments on GPT-2 small and medium models show comparable performance to baselines without recurrence, while enabling recurrence significantly reduces validation loss and improves downstream task accuracy.
This work proposes a training-free and architecture-agnostic method to enhance the inference performance of pretrained Transformers. By introducing a lightweight recurrent mechanism during inference, the approach reuses frozen intermediate contiguous layer blocks, decomposing a single large forward pass into multiple damped small-step updates. This yields the first training-free recurrent Transformer, interpreted through the lens of ordinary differential equations (ODEs) as a refined approximation of forward Euler steps. The integration of pre-normalization blocks with a damping substep strategy effectively mitigates performance degradation. Experiments demonstrate consistent gains: a 2.64% improvement on MMLU-Pro with Qwen3-4B-Instruct, a 1.14% gain on CommonsenseQA with Qwen3-30B-A3B-Instruct, and a 1.20% increase on OpenBookQA with Moonlight-16B-A3B-Instruct.
This study investigates how linear recurrent Transformers equipped with layer normalization implicitly learn the power method through gradient descent when trained on principal component prediction tasks. The work reveals an “algorithmic implicit bias”: in the absence of explicit supervision, the self-attention layers automatically converge to solutions that implement power iterations, with each layer corresponding to one update step of the power method. Theoretical analysis demonstrates that layer normalization is essential for realizing the exact power method—models without it fail to replicate the algorithm, resulting in significantly degraded performance. This paper is the first to establish the pivotal role of layer normalization in inducing algorithmic inductive biases and provides provable guarantees for the resulting performance gap.