🤖 AI Summary
To address the challenge of modeling inherent nonlinear inter-variable dependencies in time-series forecasting with Transformers, this paper proposes a novel attention mechanism based on Chatterjee’s rank correlation coefficient—the first to incorporate a nonlinear, nonparametric rank-based similarity measure into self-attention computation, replacing conventional dot-product similarity. To enable end-to-end differentiability, we introduce a differentiable SoftSort/SoftRank approximation. The mechanism integrates seamlessly into mainstream time-series Transformer architectures (e.g., Autoformer, Informer). Extensive experiments on multiple real-world datasets demonstrate consistent performance gains: average prediction error reductions of up to 9.1% over state-of-the-art (SOTA) time-series Transformers, with robust improvements across diverse forecasting horizons and data regimes. Our core contribution lies in pioneering the integration of robust, distribution-free rank correlation into attention, significantly enhancing the model’s capacity to capture complex, nonlinear temporal dependencies.
📝 Abstract
Various Transformer-based models have been proposed for time series forecasting. These models leverage the self-attention mechanism to capture long-term temporal or variate dependencies in sequences. Existing methods can be divided into two approaches: (1) reducing computational cost of attention by making the calculations sparse, and (2) reshaping the input data to aggregate temporal features. However, existing attention mechanisms may not adequately capture inherent nonlinear dependencies present in time series data, leaving room for improvement. In this study, we propose a novel attention mechanism based on Chatterjee's rank correlation coefficient, which measures nonlinear dependencies between variables. Specifically, we replace the matrix multiplication in standard attention mechanisms with this rank coefficient to measure the query-key relationship. Since computing Chatterjee's correlation coefficient involves sorting and ranking operations, we introduce a differentiable approximation employing SoftSort and SoftRank. Our proposed mechanism, ``XicorAttention,'' integrates it into several state-of-the-art Transformer models. Experimental results on real-world datasets demonstrate that incorporating nonlinear correlation into the attention improves forecasting accuracy by up to approximately 9.1% compared to existing models.