Random Feature Gaussian Process Attention: Linear-Time Probabilistic Attention with Calibrated Uncertainty

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inadequate calibration of Transformer attention mechanisms in safety-critical applications and the prohibitive computational complexity of existing Gaussian process-based approaches. To this end, we propose RFF-GPA, a plug-and-play module that leverages random Fourier features to approximate stationary kernel functions. By incorporating low-rank approximations, the proposed method reduces the computational complexity of posterior mean and variance estimation in Gaussian process attention from quadratic to linear. Experimental results demonstrate that RFF-GPA significantly improves model calibration across multiple datasets while preserving predictive accuracy. Consequently, this work effectively overcomes the sequence length scalability bottleneck inherent in probabilistic attention mechanisms, enabling their practical deployment in reliability-sensitive domains.
📝 Abstract
Transformers provide a state-of-the-art modeling framework, yet poor calibration limits their reliability in safety-critical applications. A promising direction addresses this issue by interpreting attention as a Gaussian process (GP) posterior, which enables principled uncertainty calibration but incurs cubic complexity in sequence length due to the inversion of the kernel; although decoupled GP variants reduced the cost to quadratic, the computation remains prohibitive in practice. In this paper, we propose the plug-and-play random Fourier feature Gaussian process attention (RFF-GPA) module, which represents the attention as a GP with a stationary kernel approximated by random Fourier features. This low-rank approximation results in linear-time complexity for approximating the posterior mean and variance, making it far more scalable compared to previous work. Empirical results on multiple real-world datasets show that our attention module improves calibration while maintaining predictive accuracy, and simultaneously reduces computational complexity to linear in the sequence length.
Problem

Research questions and friction points this paper is trying to address.

Transformer
uncertainty calibration
Gaussian process attention
computational complexity
scalability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Random Fourier Features
Gaussian Process Attention
Linear-Time Complexity
Uncertainty Calibration
Transformers
🔎 Similar Papers
2023-03-04International Conference on Learning RepresentationsCitations: 13
2024-03-04Computer Vision and Pattern RecognitionCitations: 3
A
Amir Mohammad Mahfoozi
Department of Computer Engineering, Sharif University of Technology
Z
Zi Yang
School of Computing and Data Science, The University of Hong Kong
Y
Ying Li
School of Computing and Data Science, The University of Hong Kong
Michael Minyi Zhang
Michael Minyi Zhang
University of Hong Kong
Bayesian non-parametricsmachine learningscalable inference