recurrent transformer

Designs and implements transformer architectures that apply the same transformer block repeatedly (looped) across depth or time with tied weights, producing recurrent or weight‑shared models that reduce stored parameters. This work covers specifying the recurrence/unrolling scheme and state/attention carryover, adapting positional and normalization components, and choosing training strategies (e.g., backpropagation through recurrence or truncated unrolling) while analyzing the resulting trade‑offs in capacity, efficiency, and convergence.

recurrenttransformer

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.42
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the training instability and poor hyperparameter transferability in Looped Transformers, which arise from highly correlated residual updates across iterations due to weight sharing. The authors demonstrate that conventional depth-scaling strategies (e.g., ε = 1/√L) are insufficient in such recurrent architectures and reveal, for the first time, that residual correlations necessitate a stronger 1/N scaling with respect to the number of loop iterations N. They propose a composite scaling strategy, ε = λ/(N√L), which decouples the learning rate from N and makes it dependent solely on the effective depth L. Grounded in theoretical analysis and multi-layer loop block modeling, experiments confirm that this approach significantly enhances training stability, consistently outperforms 1/√N scaling across varying N, and enables direct transfer of large-N models without re-tuning hyperparameters.

Hyperparameter TransferLooped TransformersResidual Scaling

This work addresses the limitations of conventional Transformers, which are constrained by finite effective depth, and recurrent models, which suffer from optimization instability and poor hardware efficiency despite their theoretically infinite depth. The authors propose a recurrent Transformer architecture that incorporates intra-layer recurrent attention, wherein each layer dynamically computes key-value pairs based on its own historical activations. This design significantly increases effective model depth while preserving the computational cost of standard autoregressive decoding. By unifying Transformer and token-level recurrence, the approach avoids training instability and achieves superior performance with fewer layers. Furthermore, a block-wise exact algorithm reduces HBM traffic from Θ(N²) to Θ(N log N), yielding an arithmetic intensity of Θ(N / log N). Experiments demonstrate that the model outperforms standard Transformer baselines of comparable size on C4 pretraining tasks with 150M and 300M parameters.

autoregressive decodinghardware efficiencyoptimization instability

Intra-Layer Recurrence in Transformers for Language Modeling

May 03, 2025
AN
Anthony Nguyen
🏛️ Algoma University

Increasing Transformer depth leads to exponential growth in parameter count, while existing recurrent methods perform coarse-grained, layer-level repetition without fine-grained control over computation. Method: We propose Intra-Layer Recurrence (ILR), a fine-grained recurrence mechanism that—within a single forward pass—selectively iterates core submodules (e.g., FFN or attention) multiple times inside a single Transformer layer, enabling dynamic state reuse without adding parameters, modifying architecture, or introducing auxiliary computational graphs. ILR employs a learnable iteration scheduling policy (e.g., allocating more iterations to earlier layers) to adaptively allocate compute resources. Contribution/Results: Evaluated on standard Transformer architectures for language modeling, ILR achieves comparable or superior performance to deeper baselines using significantly fewer parameters, thereby improving the trade-off between parameter efficiency and modeling capacity.

Optimizing layer iteration allocation for better performanceReducing parameter growth in deep transformer modelsSelectively applying recurrence within individual layers

Investigating Recurrent Transformers with Dynamic Halt

Feb 01, 2024
JR
Jishnu Ray Chowdhury
🏛️ University of Illinois at Chicago

This paper addresses the limited recursive capability of Transformers in modeling long sequences. It systematically investigates two enhancement paradigms: depth-wise recursion (e.g., Universal Transformer) and chunked temporal recursion (e.g., Temporal Latent Bottleneck). The authors propose two key innovations: (1) a dynamic halting mechanism based on global mean activation, enabling adaptive computation steps; and (2) the first integration of depth-wise recursion into a chunked temporal framework, yielding a cross-paradigm fused architecture. Extensive evaluation on the Long-Range Arena (LRA) benchmark—including ListOps, Flip-Flop, and Logical Inference—demonstrates significant differences in the effectiveness of various recursive inductive biases. Results confirm that the dynamic halting mechanism jointly improves both modeling accuracy and computational efficiency. The implementation is publicly available.

Evaluating models on diagnostic tasks like LRA and ListOpsProposing dynamic halting mechanisms for Universal TransformersStudying inductive biases in recurrent-augmented Transformer models

State-Space Modeling in Long Sequence Processing: A Survey on Recurrence in the Transformer Era

Jun 13, 2024
MT
Matteo Tiezzi
🏛️ IIT | University of Siena | IMT

Long-sequence modeling faces fundamental challenges including limited context length, difficulty in capturing long-range dependencies, and low efficiency in online learning. To address these, this work systematically reviews the resurgence of state-space models (SSMs) and recurrent computation, proposing a novel local forward-computation paradigm tailored for real-world online learning—thereby circumventing the temporal backtracking constraints inherent in standard backpropagation through time (BPTT). We introduce the first unified taxonomy encompassing both deep SSMs and large-context Transformers. Our framework integrates structured linear attention, enhanced RNN architectures, local recurrence mechanisms, and online optimization algorithms. The study rigorously clarifies the theoretical representational advantages and practical sequential reasoning benefits of recurrent modeling over alternatives. Collectively, this work delivers a scalable technical roadmap for low-latency, highly extensible long-sequence modeling.

Addressing limitations of Transformers with state-space modelsExploring efficient online learning beyond backpropagation through timeSurveying recurrent models for long sequence processing

Latest Papers

What's happening recently
View more

This study investigates under what conditions weight-shared recurrent Transformers can learn and generalize genuine algorithmic behaviors, with a focus on the group word problem. Through controlled experiments combining SGD optimization, operator-priority curriculum learning, and causal ablation cone analysis, the work reveals that models tend to learn linear sequential algorithms, with their computational frontier governed by training budget according to \( v \sim n_{\text{train}} / T_{\text{train}} \) (R² = 0.99). The findings indicate that architectural priors dominate algorithm selection, and group order—not circuit complexity—constitutes the primary learning bottleneck. The paper proposes a convergence-rate-based stopping rule, a transferable mechanism, and introduces a novel head-measurement tool \( \tau(n,i) \). Full generalization is achieved on the A5 group, and results are successfully replicated on an easy-to-hard benchmark.

algorithmic convergencecomputation frontiergroup word problems

This work addresses the instability in recurrent Transformers caused by parameter sharing, particularly the mismatch between forward propagation and gradient aggregation when a module is accessed multiple times. To resolve this, the study introduces the access-aligned coefficient κ_R, which explicitly incorporates the number of parameter accesses into residual scaling design to model parameter update dynamics under recurrent depth. Through first-order perturbation analysis, the authors demonstrate that stable training requires increasing the scaling exponent in Post-LN DeepNorm from 1/4 to 1/2, leading to the proposed DeepLoop method. Experiments on GPT-2 small and medium models show comparable performance to baselines without recurrence, while enabling recurrence significantly reduces validation loss and improves downstream task accuracy.

depth scalingLooped Transformersparameter sharing

This work proposes a training-free and architecture-agnostic method to enhance the inference performance of pretrained Transformers. By introducing a lightweight recurrent mechanism during inference, the approach reuses frozen intermediate contiguous layer blocks, decomposing a single large forward pass into multiple damped small-step updates. This yields the first training-free recurrent Transformer, interpreted through the lens of ordinary differential equations (ODEs) as a refined approximation of forward Euler steps. The integration of pre-normalization blocks with a damping substep strategy effectively mitigates performance degradation. Experiments demonstrate consistent gains: a 2.64% improvement on MMLU-Pro with Qwen3-4B-Instruct, a 1.14% gain on CommonsenseQA with Qwen3-30B-A3B-Instruct, and a 1.20% increase on OpenBookQA with Moonlight-16B-A3B-Instruct.

inference-timelooped transformersperformance improvement

This study investigates how linear recurrent Transformers equipped with layer normalization implicitly learn the power method through gradient descent when trained on principal component prediction tasks. The work reveals an “algorithmic implicit bias”: in the absence of explicit supervision, the self-attention layers automatically converge to solutions that implement power iterations, with each layer corresponding to one update step of the power method. Theoretical analysis demonstrates that layer normalization is essential for realizing the exact power method—models without it fail to replicate the algorithm, resulting in significantly degraded performance. This paper is the first to establish the pivotal role of layer normalization in inducing algorithmic inductive biases and provides provable guarantees for the resulting performance gap.

Algorithmic Implicit BiasLayer NormalizationPower Method

Hot Scholars

YK

Yoon Kim

Associate Professor, MIT
Machine LearningNatural Language ProcessingDeep Learning
YI

Yusuke Iwasawa

The University of Tokyo
deep learningtransfer learningfoundation modelmeta learning
KH

Kohei Hayashi

Researcher, Preferred Networks
Machine LearningWeb Data MiningTensor Decomposition
YM

Yutaka Matsuo

Department of Physics, Graduate School of Science, The University of Tokyo
String theoryMathematical PhysicsQuantum Field Theory