🤖 AI Summary
This study elucidates the control mechanisms by which feed-forward networks (FFNs) govern token dynamics within Transformers. Adopting a cybernetic perspective, we model self-attention layers as interacting particle systems and treat FFNs as independent controllers. Theoretical analysis demonstrates that FFNs can arbitrarily steer tokens toward consensus or clustering, entirely decoupled from QKV parameters. Through numerical simulations and empirical comparisons with large language models (LLMs), we validate the consensus-convergence capability of FFNs and reveal a strong alignment between theoretical predictions and the threshold behaviors observed in real-world LLMs. This work establishes a novel paradigm for understanding the internal dynamics of Transformer architectures.
📝 Abstract
We study the dynamical behavior of tokens in transformers from a control-theoretic perspective. Our model includes the feed-forward layer present after the self-attention mechanism, with the self-attention mechanism interpreted as an interacting particle system and the feed-forward layer as an independent control. Our main theoretical result establishes that the feed-forward network can steer the tokens arbitrarily close to consensus regardless of the key, query, and value matrices. Our result are easily extended to convergence to many clusters and to multi-head attention. We conduct numerical experiments to verify our results, and compare thresholding behavior from our theory to real-world LLMs.