XicorAttention: Time Series Transformer Using Attention with Nonlinear Correlation

📅 2025-06-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the challenge of modeling inherent nonlinear inter-variable dependencies in time-series forecasting with Transformers, this paper proposes a novel attention mechanism based on Chatterjee’s rank correlation coefficient—the first to incorporate a nonlinear, nonparametric rank-based similarity measure into self-attention computation, replacing conventional dot-product similarity. To enable end-to-end differentiability, we introduce a differentiable SoftSort/SoftRank approximation. The mechanism integrates seamlessly into mainstream time-series Transformer architectures (e.g., Autoformer, Informer). Extensive experiments on multiple real-world datasets demonstrate consistent performance gains: average prediction error reductions of up to 9.1% over state-of-the-art (SOTA) time-series Transformers, with robust improvements across diverse forecasting horizons and data regimes. Our core contribution lies in pioneering the integration of robust, distribution-free rank correlation into attention, significantly enhancing the model’s capacity to capture complex, nonlinear temporal dependencies.

Technology Category

Machine Learning: Learning Preferences or RankingsComputer Vision: Diffusion Models for VisionReasoning under Uncertainty: Relational Probabilistic Models

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingGraph Algorithms and Modeling for the Web: Algorithms and analysis for heterogeneous, signed, attributed, multi-relational, temporal, higher-order, and annotated Web-related graphs
📝 Abstract
Various Transformer-based models have been proposed for time series forecasting. These models leverage the self-attention mechanism to capture long-term temporal or variate dependencies in sequences. Existing methods can be divided into two approaches: (1) reducing computational cost of attention by making the calculations sparse, and (2) reshaping the input data to aggregate temporal features. However, existing attention mechanisms may not adequately capture inherent nonlinear dependencies present in time series data, leaving room for improvement. In this study, we propose a novel attention mechanism based on Chatterjee's rank correlation coefficient, which measures nonlinear dependencies between variables. Specifically, we replace the matrix multiplication in standard attention mechanisms with this rank coefficient to measure the query-key relationship. Since computing Chatterjee's correlation coefficient involves sorting and ranking operations, we introduce a differentiable approximation employing SoftSort and SoftRank. Our proposed mechanism, ``XicorAttention,'' integrates it into several state-of-the-art Transformer models. Experimental results on real-world datasets demonstrate that incorporating nonlinear correlation into the attention improves forecasting accuracy by up to approximately 9.1% compared to existing models.
Problem

Research questions and friction points this paper is trying to address.

Capturing nonlinear dependencies in time series data
Improving Transformer attention with rank correlation
Enhancing forecasting accuracy via nonlinear correlation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Novel attention mechanism using Chatterjee's rank correlation
Differentiable approximation with SoftSort and SoftRank
Improves forecasting accuracy by 9.1%