🤖 AI Summary
This study addresses the observation that representational similarity within the residual stream of Transformers decays with increasing layer distance, a structural property for which effective quantification tools remain lacking. By modeling the residual stream as a discrete-time process, this work proposes an "effective depth" metric that diagnoses the global cumulative state of a model through aggregated similarity decay curves. The analysis reveals that correlations among residual updates cause empirical measurements to fall below theoretical upper bounds. Experiments across 16 language models validate the proposed metric, demonstrating its capacity to distinguish structural consequences from empirical phenomena while remaining robust against common artifacts. These findings establish effective depth as a reliable diagnostic tool for analyzing Transformer architectures.
📝 Abstract
We introduce \emph{effective depth} ($\Deff$), a scalar diagnostic that treats the layer-wise residual stream of a transformer as a discrete-time process, measures how representation similarity decays with layer distance, and aggregates that profile into one number. Across sixteen decoder-only language models, $\Deff$ separates a structural consequence of residual accumulation from an empirical one: even maximally diverse orthogonal updates have the closed-form reference $F_L = 2L/(L+1)<2$, yet fifteen of sixteen default measurements lie below $F_L$ (Qwen3.5: 32--44\%, OLMo-2: 40--41\%, Pythia: 23--28\%). Matched references show that the gap is not caused by the persistent initial state or update-size imbalance, but is largely a calibrated signature of correlated residual updates rather than evidence that depth is unused. Symmetric position-0, token-normalisation, and top-PC controls show the regime is not reducible to BOS or top-PC artefacts: the lone above-reference default outlier joins the same regime, and all sixteen models are sub-reference after token-normalisation or top-1-PC removal. Intermediate checkpoints show that the regime is established early in OLMo-2 and stable through 5T tokens, while Pythia-1.4B follows a distinct decreasing trajectory. A controlled residual-carry intervention supports the mechanism, and $\Deff$ is best read as a \emph{global} accumulated-state diagnostic, not as a capability score or pruning method.