🤖 AI Summary
This study addresses the inability of existing theories to explain the rich structural patterns in Transformer token representations. For the first time, pattern formation theory is introduced into the complete Transformer architecture, unifying positional encoding, multi-head attention, and output geometry within a dynamical systems framework. This perspective elucidates how individual components regulate pattern amplification and stabilization, revealing their underlying inductive biases. Furthermore, the research identifies and characterizes previously unreported dynamical patterns, including traveling and rotating waves, along with their mechanisms of competition and coexistence. Empirical evaluations demonstrate that incorporating controllable dynamical priors significantly enhances learning efficiency, yielding notable improvements in data efficiency and optimization speed across sequence modeling tasks and ConViT on CIFAR-10.
📝 Abstract
What are the inductive biases of a Transformer architecture? Existing theory on how the forward pass shapes representations either considers whether Transformers escape from rank collapse or demonstrates that self-attention drives tokens toward cluster patterns. The latter view arises from an elegant dynamical systems perspective, but relies on simplified architectural assumptions, and does not explain the rich structures observed in practice. This leaves a major open question: when a full Transformer escapes rank collapse, how does it structure token representations? Using pattern-formation theory, we show that the dynamical view of Transformers can account for Positional Encoding, Multi-Head Attention, and Output-Value geometry. We demonstrate that a full Transformer architecture imposes an inductive prior by selectively amplifying a rich set of previously unreported patterns, including traveling or rotating waves among others. We characterize the role of each architectural component in controlling which pattern is amplified, which ones stabilize, compete, or coexist. Finally, we show that these structures can act as a controllable dynamical prior that facilitates learning. By choosing both task-aligned positional encoding and weight initialization, we demonstrate improved data efficiency and accelerated optimization on controlled sequence tasks and with ConViT on CIFAR-10.