🤖 AI Summary
This study addresses the limitation of existing research that treats attention heads as discrete dictionary entries, lacking a unified perspective on the contextualized global behavior of language models. We propose the principle that low-information words absorb more context, reinterpreting attention heads as continuous channels to optimize information routing. By leveraging singular vectors to direct enhanced writing toward high-information words, this work achieves continuous-space embedding and automated interpretation of attention heads. Combining singular value decomposition with hidden state analysis, our approach reveals general principles underlying model contextualization, improves information transmission precision, and enables automated interpretation of attention heads that surpasses conventional heuristic methods.
📝 Abstract
Contextualization, the core operation of language modeling, transmits information across words to build sentence-specific word representations. Prior works mainly study contextualization, focusing on individual words and attention heads as a growing discrete dictionary, lacking a global view of their general behavior. Therefore, we propose a general principle: Globally, we find and estimate that different words carry different amounts of information, and less-informative words tend to absorb more contextual information. Specifically, these low-information words do not absorb contextual words uniformly, and finer-grained selectivity enables more precise routing to promote information transmission between matched words. Moreover, to find what mechanism causes such processing, we reinterpret attention heads as channels gated by their singular vectors and find that: (1) these singular vectors point to the hidden states of more informative words, allowing such words to write their information to others more strongly to act as information sources, and vice versa; and (2) these singular vectors can be viewed equally as hidden state features, enabling automated interpretation of attention heads beyond prior heuristic head discovery, also embedding heads into a continuous space rather than treating them as discrete, independent dictionary entries.