The Residual Stream's Effective Depth

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the observation that representational similarity within the residual stream of Transformers decays with increasing layer distance, a structural property for which effective quantification tools remain lacking. By modeling the residual stream as a discrete-time process, this work proposes an "effective depth" metric that diagnoses the global cumulative state of a model through aggregated similarity decay curves. The analysis reveals that correlations among residual updates cause empirical measurements to fall below theoretical upper bounds. Experiments across 16 language models validate the proposed metric, demonstrating its capacity to distinguish structural consequences from empirical phenomena while remaining robust against common artifacts. These findings establish effective depth as a reliable diagnostic tool for analyzing Transformer architectures.
📝 Abstract
We introduce \emph{effective depth} ($\Deff$), a scalar diagnostic that treats the layer-wise residual stream of a transformer as a discrete-time process, measures how representation similarity decays with layer distance, and aggregates that profile into one number. Across sixteen decoder-only language models, $\Deff$ separates a structural consequence of residual accumulation from an empirical one: even maximally diverse orthogonal updates have the closed-form reference $F_L = 2L/(L+1)<2$, yet fifteen of sixteen default measurements lie below $F_L$ (Qwen3.5: 32--44\%, OLMo-2: 40--41\%, Pythia: 23--28\%). Matched references show that the gap is not caused by the persistent initial state or update-size imbalance, but is largely a calibrated signature of correlated residual updates rather than evidence that depth is unused. Symmetric position-0, token-normalisation, and top-PC controls show the regime is not reducible to BOS or top-PC artefacts: the lone above-reference default outlier joins the same regime, and all sixteen models are sub-reference after token-normalisation or top-1-PC removal. Intermediate checkpoints show that the regime is established early in OLMo-2 and stable through 5T tokens, while Pythia-1.4B follows a distinct decreasing trajectory. A controlled residual-carry intervention supports the mechanism, and $\Deff$ is best read as a \emph{global} accumulated-state diagnostic, not as a capability score or pruning method.
Problem

Research questions and friction points this paper is trying to address.

effective depth
residual stream
transformer
representation similarity
residual accumulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

effective depth
residual stream
transformer diagnostics
representation similarity
discrete-time process
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Barak Gahtan
Technion Israel Institute of Technology
Ido Galil
Ido Galil
Technion - Israel Institute of Technology
Machine learningDeep learning
A
Alex M. Bronstein
Technion Israel Institute of Technology & ISTA Institute of Science and Technology Austria