LayerRoPE: Dynamic Depth-wise Magnitude & Angular Superposition

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the "curse of depth" in Transformers, where hidden state norms explode in deep layers. We propose LayerRoPE, a method that departs from conventional paradigms suppressing residual streams. Inspired by Rotary Position Embedding (RoPE), it reformulates layer normalization weights as depth positional encodings, jointly encoding hierarchical information through the magnitude and phase of a shared γ vector to enable depth-conditioned modulation with substantially fewer parameters. Experiments demonstrate that LayerRoPE achieves stable convergence at 512 layers, matching Pre-Norm performance with 3.4× less computation while improving learning rate tolerance by 3–10×. Furthermore, the approach transfers seamlessly to hybrid architectures and recurrent latent models.
📝 Abstract
As data propagates through a Transformer, the norm of its hidden states grows by orders of magnitude with depth, a phenomenon framed as 'curse of depth' and nearly universally treated as a pathology to be suppressed. We take the opposite view. Across 16 pre-trained LLMs from 9 families, spanning dense, mixture-of-experts and hybrid architectures and Pre-, Peri- and Post-Norm designs, we find that this growth reflects an emergent depth-positional encoding, carried by the only learned per-layer gain on the residual stream, the normalization weight $γ$: with depth, $γ$ grows in magnitude and rotates in direction, jointly encoding the layer index. We make this depth-conditioned encoding explicit with LayerRoPE, an implicit analog of RoPE along the depth axis, which replaces all layerwise $γ$ vectors with a single shared vector and depth-conditioned scalars, at a net reduction in parameters and $<0.02\%$ change in FLOPs. Across a model ladder scaled up to $100$B+ tokens, LayerRoPE consistently outperforms Pre-, Post- and Peri-Norm and Layer-Norm Scaling, reaching Pre-Norm's 1.3B loss with $3.4\times$ less compute; LayerRoPE is the only approach that shows strong convergence and improves near monotonically as depth scales to 512 layers. It improves learning-rate sensitivity by $3$-$10\times$, and transfers naively to and consistently improves looped latent models and Vision Transformers. Inspecting its learned schedule inverts the prevailing premise: LayerRoPE does not shrink the residual stream but widens it, damping what each block reads while amplifying what it writes. Depth stability, our results suggest, calls not for suppressing the residual stream, but for depth-conditioned regulation of the computational blocks it feeds.
Innovation

Methods, ideas, or system contributions that make the work stand out.

LayerRoPE
depth-positional encoding
residual stream regulation
normalization weight
depth stability
🔎 Similar Papers
No similar papers found.