🤖 AI Summary
This study addresses the inadequate calibration of Transformer attention mechanisms in safety-critical applications and the prohibitive computational complexity of existing Gaussian process-based approaches. To this end, we propose RFF-GPA, a plug-and-play module that leverages random Fourier features to approximate stationary kernel functions. By incorporating low-rank approximations, the proposed method reduces the computational complexity of posterior mean and variance estimation in Gaussian process attention from quadratic to linear. Experimental results demonstrate that RFF-GPA significantly improves model calibration across multiple datasets while preserving predictive accuracy. Consequently, this work effectively overcomes the sequence length scalability bottleneck inherent in probabilistic attention mechanisms, enabling their practical deployment in reliability-sensitive domains.
📝 Abstract
Transformers provide a state-of-the-art modeling framework, yet poor calibration limits their reliability in safety-critical applications. A promising direction addresses this issue by interpreting attention as a Gaussian process (GP) posterior, which enables principled uncertainty calibration but incurs cubic complexity in sequence length due to the inversion of the kernel; although decoupled GP variants reduced the cost to quadratic, the computation remains prohibitive in practice. In this paper, we propose the plug-and-play random Fourier feature Gaussian process attention (RFF-GPA) module, which represents the attention as a GP with a stationary kernel approximated by random Fourier features. This low-rank approximation results in linear-time complexity for approximating the posterior mean and variance, making it far more scalable compared to previous work. Empirical results on multiple real-world datasets show that our attention module improves calibration while maintaining predictive accuracy, and simultaneously reduces computational complexity to linear in the sequence length.