🤖 AI Summary
This study addresses the inherent complexity of Transformer internal mechanisms, which impedes intuitive understanding of how geometric representations translate into algorithmic logic. We construct a minimal Transformer with an embedding dimension of two and propose a "geometry as algorithm" framework. Through comprehensive two-dimensional visualization of attention patterns, residual streams, and decision boundaries, this approach directly interprets learned geometric structures as stepwise algorithmic logic. The proposed method enables explicit visualization and interpretability analysis of the model's entire internal computation pipeline. Furthermore, we successfully reproduce and dissect the predictive mechanisms on numerical sequence tasks. Ultimately, this work provides an innovative pedagogical and experimental platform for advancing mechanistic interpretability research in deep learning.
📝 Abstract
We present a framework for building and interpreting minimal transformer models. By constraining a transformer's embedding dimension and head size to 2, we enable full two-dimensional visualization of its internal representations. Embeddings, query/key/value transforms, attention outputs, residual streams, and decision boundaries can all be seen directly. Our central claim is that the learned geometry implies an algorithm; the arrangement of points and boundaries in R^2 can be read as a step-by-step procedure. We train a transformer on a simple task where it must produce the most recently observed even number whenever the '+' operator appears in a sequence of digits. Once trained, we visually walk through every step of the transformer's computation. We show how the model embeds the tokens and their respective positions in the sequence, transforms them via the Q, K, and V matrices, uses the dot product between the Q and K representations to form the attention matrix, and uses the attention matrix to select values that move the representation of each input token to the region of the domain of the output layer that will correctly predict the next token. We introduce a suite of interpretability visualizations that make the algorithmic interpretation of this procedure explicit. Our framework offers a pedagogical and experimental testbed to explore how transformers use informational geometry to implement next-token prediction.