π€ AI Summary
This study investigates whether deep self-attention models with limited parameter sharing can achieve universal approximation solely by increasing depth. To this end, it proposes a "dual" mechanism that eliminates the need for infinite width, enabling universal interpolation mappings between arbitrary sequences through a fixed set of attention blocks and their prescribed composition order. Technically, the approach employs Gaussian-initialized projection matrices, residual softmax attention, and causal masking. Theoretically, this work demonstrates that only two frozen residual blocks suffice to attain universality, establishing universal interpolation guarantees under continuous and finite-depth settings while elucidating the impact of causal constraints. These findings provide rigorous theoretical foundations for understanding the expressive power of deep Transformers.
π Abstract
Universal approximation is a necessary qualitative property of learning architectures to benefit from scaling laws. While it is generically verified on a variety of neural architectures and random feature models, it typically involves infinite width limits. In this work, we focus on deep self-attention models and consider instead the `dual' regime, where approximation power is enabled entirely by depth, and featuring strong parameter sharing across layers, motivated by recent models such as the Looped Transformers. More specifically, we ask whether one can find a predefined finite set of parameters, each defining an attention block, such that the resulting finite set of transformations can map any collection of $N$ sequences of $n$ tokens to any other collection of $N$ sequences of $n$ tokens. Crucially, these transformations are \emph{fixed independently of the input and output} collections: only the order in which the blocks are applied, their signs, and their durations depend on the particular interpolation task. Our main result establishes it for residual softmax attention using only two frozen single-head blocks with Gaussian-initialized projection matrices. The result holds at both continuous and finite depth. We also characterize the restrictions imposed by causal masking and establish corresponding universal interpolation guarantees.