🤖 AI Summary
This work addresses the “lost-in-the-middle” phenomenon in Transformer decoders—where information located in the middle of long prompts is poorly retrieved—by providing, for the first time, a rigorous theoretical explanation from a dynamical systems perspective. By modeling causal self-attention as a non-commutative interacting particle system and leveraging cumulant expansions under triangular causal structure together with Glauber calculus, the authors establish a mean-field limit and analytically solve the associated correlation equations. Under the assumption of i.i.d. uniformly distributed inputs, they rigorously prove that token retrieval performance exhibits a U-shaped dependence on source position, thereby quantitatively elucidating the origins of primacy and recency effects alongside the characteristic performance dip in the middle of the sequence.
📝 Abstract
We study causal self-attention dynamics -- a toy model for decoder Transformers -- which we interpret as a non-exchangeable interacting particle system. Adapting cumulant expansions to the triangular causal dependency structure of the model, and appealing to non-hierarchical methods to estimate correlations using Glauber calculus, we prove a quantitative mean-field limit result and a next-order characterization of correlations. For iid uniformly distributed tokens, the limiting correlation equation can be solved in closed form and we obtain a rigorous explanation of the empirically observed \emph{lost-in-the-middle} phenomenon: the token retrieval profile, as a function of the source position in the prompt, is $\mathsf{U}$-shaped, with primacy, recency, and a unique interior minimum under an explicit smallness condition.