🤖 AI Summary
This work addresses the limitation of conventional Transformers, which employ fixed additive residual connections along the depth axis and thus lack adaptive inter-layer information aggregation. The authors propose a dual perspective on residual flows, conceptualizing the decoder as an evolving system across two ordered dimensions—sequence positions and layer depth—and establish, for the first time, an operator-level equivalence between causal residual attention along the depth axis and short sliding window attention (ShortSWA) along the sequence axis. This duality provides a unified interpretation of methods such as ELC-BERT, DenseFormer, and Vertical Attention, while distinguishing operator-level duality from system-level asymmetry. Building on this insight, the study integrates Deep Delta Learning (DDL) with ShortSWA to demonstrate that, in large-scale autoregressive models, ShortSWA aligns better with hardware characteristics, whereas DDL more effectively optimizes residual pathways, thereby offering clear principles for architectural design.
📝 Abstract
Recent work has made clear that the residual pathway is not mere optimization plumbing; it is part of the model's representational machinery. We agree, but argue that the cleanest way to organize this design space is through a two-axis view of the Transformer. A decoder evolves information along two ordered dimensions: sequence position and layer depth. Self-attention already provides adaptive mixing along the sequence axis, whereas the residual stream usually performs fixed addition along the depth axis. If we fix a token position and treat layer index as the ordered variable, then a causal depth-wise residual attention read is exactly the same local operator as causal short sliding-window attention (ShortSWA), except written over depth rather than over sequence. This is the core residual stream duality behind Transformer$^2$. This perspective also clarifies the recent literature. ELC-BERT and DenseFormer already show that learned aggregation over depth can outperform uniform residual accumulation, while Vertical Attention, DeepCrossAttention (DCA), MUDDFormer, and Attention Residuals move further toward explicit attention-based routing over earlier layers. The key point, however, is that operator-level duality does not imply systems-level symmetry. For large-scale autoregressive models, sequence-axis ShortSWA is usually the more hardware-friendly placement because it reuses token-side sliding-window kernels, KV-cache layouts, and chunked execution. If the goal is instead to change the shortcut itself, Deep Delta Learning (DDL) is the cleaner intervention because it modifies the residual operator directly rather than adding a separate cross-layer retrieval path. Our recommendation is therefore simple: use DDL when the shortcut is the object of interest, and use sequence-axis ShortSWA when the goal is local adaptive mixing.