Attention Kernels for Learning Maps Between Heavy-Tailed Measures

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the integral divergence and model collapse issues arising from exponential attention in Transformers under heavy-tailed distributions. We propose replacing standard softmax with slow-growing kernels and introducing Symlog data preprocessing. By constructing a dedicated benchmark and employing Wasserstein distance metrics, we systematically evaluate the synergistic effects of slow-growing kernels and data transformations on operator learning within post-normalization architectures. Experimental results demonstrate that the proposed method effectively overcomes integral collapse, significantly outperforming conventional softmax attention mechanisms. This work establishes a theoretically grounded and practically valuable new paradigm for operator learning on heavy-tailed data.
📝 Abstract
Operator learning on probability measures can be accomplished with transformers. For measures with polynomial tails, the exponential weighting in softmax can make the corresponding measure-level attention integrals diverge. This motivates replacing the exponential with slower-growing functions. We construct two benchmarks for operator learning on measures with closed-form targets. We use these benchmarks to study attention kernel growth and data transformation in post-norm transformers. Without data transformation, the softmax models exhibit ensemble collapse on both heavy-tailed benchmarks, while the three slower-growing kernels avoid collapse. Symlog preprocessing allows softmax to avoid collapse on the matrix inverse task but not on the sheared swap task. On the Gaussian control, all four kernels perform similarly. We also examine how sample size affects the sensitivity of empirical energy and Wasserstein distances to tail differences. These results support slower-growing attention kernels as an effective design choice for post-norm transformers learning from heavy-tailed ensembles.
Problem

Research questions and friction points this paper is trying to address.

operator learning
heavy-tailed measures
attention kernels
softmax divergence
ensemble collapse
Innovation

Methods, ideas, or system contributions that make the work stand out.

attention kernels
heavy-tailed measures
operator learning
post-norm transformers
ensemble collapse
K
Kailen Hargenrader
Computing and Mathematical Sciences, California Institute of Technology
E
Edoardo Calvello
Lawrence Berkeley National Laboratory, University of California, Berkeley, ICSI
Bohan Chen
Bohan Chen
University of Liverpool
Artificial IntelligenceGenerative Models