π€ AI Summary
This study addresses the challenges of opaque information routing and inaccurate uncertainty estimation in Transformer attention mechanisms. Grounded in spectral geometry and operator theory, this work models attention heads as functional mappings within a Hilbert space and derives token difference operators to reveal intrinsic routing structures. By constructing a token difference spectrum under probabilistic geometry, it effectively decouples the sink effect from genuine routing capacity, providing a unified explanation for attention collapse. Ultimately, this project establishes a systematic analytical framework for attention mechanisms and introduces a novel uncertainty estimator that significantly improves both estimation accuracy and model performance on long-context tasks.
π Abstract
In this work, we study transformer attention through the lens of spectral geometry and operator theory. We view each attention head as a functional map between Hilbert spaces of functions on the token sequence and derive a Token Difference Operator, whose spectral structure controls how token-space information is routed to the output. We show that standard Euclidean spectra are structurally biased by sinks, conflating mass concentration with genuine routing capacity. By recasting token space in the intrinsic probability geometry induced by attention, the token difference spectrum disentangles sink effects from routing capacity and provides a spectral description of the dimensionality of the head output. This yields a unified framework for analyzing attention maps, explaining sinks, routing collapse, and output dimensionality within a single operator-theoretic framework. In practice, by grounding attention heuristics in spectral geometry, we develop a novel attention-based uncertainty estimator that complements probability-based scores, with the largest gains on long-context inputs.