Multiscale POD of Transformer Attention Fields: Scale-Selective Analysis via Morlet Scalogram

πŸ“… 2026-06-04
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study investigates how to extract scale-selective dominant modes from Transformer attention fields and quantify inter-layer complexity. To this end, it introduces Proper Orthogonal Decomposition (POD)β€”a technique from turbulence analysisβ€”into the study of attention mechanisms for the first time, combining it with the Morlet continuous wavelet transform to construct a multi-scale analysis framework that requires neither architectural modifications nor linguistic annotations. The work further proposes a spectral concentration index to measure the complexity of individual layers. Experimental results reveal a systematic scale organization within Transformers: shallow layers emphasize fine-grained patterns, while deeper layers shift toward coarse-grained structures. Additionally, the framework provides data-driven effective rank estimates for each layer, significantly enhancing the understanding of the intrinsic structure of attention mechanisms.
πŸ“ Abstract
We introduce scale-selective Proper Orthogonal Decomposition (POD) for transformer attention fields, inspired by the use of POD for extracting energetically dominant modes from turbulent flow ensembles. The Morlet continuous wavelet transform identifies dominant temporal scales in the attention lag structure across a document ensemble; POD then extracts the energetically dominant modes at each scale from the ensemble of attention fields. The resulting modes reveal layer-dependent scale organisation, with early layers emphasising fine scales and later layers shifting toward coarser scales. We define a spectral concentration index from the POD eigenvalue decay rate and show empirically that it differentiates layers by their attention field complexity. By the classical POD optimality theorem, the extracted modes minimise the average L2 reconstruction error over the ensemble (Theorem 1), giving a data-driven effective rank for each layer. The method requires no architectural modification and no linguistic annotations: dominant attention patterns emerge from ensemble statistics alone. The turbulence analogy is structural rather than physical: we borrow ensemble covariance and modal analysis, not fluid dynamics itself.
Problem

Research questions and friction points this paper is trying to address.

Transformer attention fields
multiscale analysis
Proper Orthogonal Decomposition
Morlet wavelet
scale-selective
Innovation

Methods, ideas, or system contributions that make the work stand out.

scale-selective POD
Morlet scalogram
attention fields
ensemble modal analysis
spectral concentration index
A
Athanasios Zeris
Independent Researcher, Athens, Greece