🤖 AI Summary
This study addresses the integral divergence and model collapse issues arising from exponential attention in Transformers under heavy-tailed distributions. We propose replacing standard softmax with slow-growing kernels and introducing Symlog data preprocessing. By constructing a dedicated benchmark and employing Wasserstein distance metrics, we systematically evaluate the synergistic effects of slow-growing kernels and data transformations on operator learning within post-normalization architectures. Experimental results demonstrate that the proposed method effectively overcomes integral collapse, significantly outperforming conventional softmax attention mechanisms. This work establishes a theoretically grounded and practically valuable new paradigm for operator learning on heavy-tailed data.
📝 Abstract
Operator learning on probability measures can be accomplished with transformers. For measures with polynomial tails, the exponential weighting in softmax can make the corresponding measure-level attention integrals diverge. This motivates replacing the exponential with slower-growing functions. We construct two benchmarks for operator learning on measures with closed-form targets. We use these benchmarks to study attention kernel growth and data transformation in post-norm transformers. Without data transformation, the softmax models exhibit ensemble collapse on both heavy-tailed benchmarks, while the three slower-growing kernels avoid collapse. Symlog preprocessing allows softmax to avoid collapse on the matrix inverse task but not on the sheared swap task. On the Gaussian control, all four kernels perform similarly. We also examine how sample size affects the sensitivity of empirical energy and Wasserstein distances to tail differences. These results support slower-growing attention kernels as an effective design choice for post-norm transformers learning from heavy-tailed ensembles.