Which the Eye Fears: Writing with Read-Blindness Explains Massive Activations in Transformers

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates why massive activations persist in Transformers even when suppression capabilities are present. Through operator-level mechanistic analysis, gradient inspection, and training checkpoint tracing, the authors reveal a read-write asymmetry between attention and feed-forward layers. The key contribution is the discovery that "reading blind spots" emerge prior to feed-forward amplification, demonstrating that they constitute an upstream mechanism actively maintained by the model. Experiments show that removing local reading masks triggers compensatory shifts, yet massive activations remain stubbornly persistent, uncovering their deep structural origins.
📝 Abstract
Massive activation features (MAs) in Transformers are extreme-value residual-stream features that persist across layers despite the model's ability to suppress them. Why do they survive? Our investigation using an operator-level mechanistic analysis of attention and feed-forward (FFN) blocks reveals that these blocks systematically ignore MA coordinates while reading, but not while writing; creating a read-write asymmetry that blocks corrective feedback while allowing continued accumulation. We find that both attention and feed-forward layers have this read-blindness, and contribute to the emergence and persistence of MAs. To validate prior work that hypothesized that FFN's amplification abilities is the primary reason for MAs (Sun et al., 2026), we analyze the model checkpoints during learning. Contrary to our expectation, read-blindness emerges before FFN amplification, suggesting that it acts upstream in the MA mechanism. We further contribute gradient analysis to link this behavior to surprising asymmetries in the loss landscape, concluding that the model actively maintains this read-blindness. Finally, we find that removing read-blocking at different locations induces compensatory shifts elsewhere, but MAs still persist.
Problem

Research questions and friction points this paper is trying to address.

Massive Activations
Transformers
Read-Blindness
Residual Stream
Mechanistic Interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Massive Activations
Read-Blindness
Read-Write Asymmetry
Mechanistic Interpretability
Transformers
🔎 Similar Papers
2023-12-17Bulletin of the American Mathematical SocietyCitations: 59