Query Expansion and Key Specialization in Transformer Attention Geometry

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the divergent geometric evolution of query (Q) and key (K) projections during Transformer training and its implications for the attention mechanism. By systematically monitoring the training trajectories of small GPT models through participation ratio analysis, effective dimensionality tracking, and spectral control experiments, this work reveals, for the first time, the dynamic evolution of Q/K geometric asymmetry: the effective dimensionality of Q expands while that of K contracts. Furthermore, it establishes a causal link between this asymmetry and changes in attention entropy. Experimental results confirm that spectral contraction in K directly drives the sharpening of attention distributions. The findings also characterize the universality of these geometric trends during early training and their subsequent decay in later stages, offering a novel perspective for understanding attention mechanisms.
📝 Abstract
The projection of queries and keys are central to the attention mechanism in Transformer architectures. While they are mathematically symmetric, they play different roles in attention mechanisms. The question of whether there is an effect from their functional distinction on their geometric development in training remains unanswered. We investigate the problem through the training of small GPT-like Transformers on character-level WikiText-103 for three different depths (4, 6, and 8 layers), three types of initialization for queries and keys, and four random seeds, resulting in 36 runs and 54 trajectories of average layers across seeds. We track the effective dimensionality of those layers using participation ratios and discover that effective dimension of queries expand while keys shrink, and that $PR_Q - PR_K$ is positive in all trajectories studied. In connection to attention, the shrinking of keys leads to a narrower spectrum of $QK^\top$ and more peaked attention weights. In order to determine if this connection is causal or coincidental, we directly control the spectrum of keys during training across five seeds: restricting it to make it shrink sharpens the attention with high directional confidence, while keeping it constant to the level of initial dispersion makes attention softer. Additional token-level checkpoint analyses show that the monotonic paired-contrast trend is not universal across pretrained families, but survives as an early-training regime that later decays over a full pretraining run, and the link between interaction-rank geometry and attention entropy remains visible in several models.
Problem

Research questions and friction points this paper is trying to address.

Transformer
attention mechanism
query-key asymmetry
effective dimensionality
geometric evolution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Query-Key Asymmetry
Effective Dimensionality
Attention Geometry
Participation Ratio
Spectrum Control
🔎 Similar Papers
V
Vidit Gupta
Dwarkadas J. Sanghvi College of Engineering, Mumbai, India
S
Siddhesh Nadkarni
Dwarkadas J. Sanghvi College of Engineering, Mumbai, India
M
Mihik Chaudhari
Dwarkadas J. Sanghvi College of Engineering, Mumbai, India
V
Vinaya Sawant
Dwarkadas J. Sanghvi College of Engineering, Mumbai, India
P
Prachi Tawde
Dwarkadas J. Sanghvi College of Engineering, Mumbai, India