Sequence Complementor: Complementing Transformers For Time Series Forecasting with Learnable Sequences

๐Ÿ“… 2025-01-06
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Transformers suffer from inadequate sequential representation learning in time-series forecasting, struggling to model multivariate channel correlations and handle data instability. Method: This paper proposes Learnable Sequence Complementorsโ€”a novel mechanism that augments input sequences to enhance representational diversity. Grounded in an information-theoretic analysis, it first establishes a linear relationship between sequence diversity (measured by entropy) and prediction error. Leveraging this insight, we design a provably sound complement attention mechanism and introduce a theoretically grounded diversity loss, jointly optimized with entropy regularization. Contribution/Results: Extensive experiments demonstrate that our approach consistently outperforms state-of-the-art methods on both long- and short-term forecasting benchmarks, achieving significant reductions in MSE. The results empirically validate that increasing sequential representation diversity is critical for improving forecasting accuracy.

Technology Category

Machine Learning: Time-Series/Data StreamsSearch and Optimization: Learning to SearchPlanning, Routing, and Scheduling: Learning for Planning and Scheduling

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphs
๐Ÿ“ Abstract
Since its introduction, the transformer has shifted the development trajectory away from traditional models (e.g., RNN, MLP) in time series forecasting, which is attributed to its ability to capture global dependencies within temporal tokens. Follow-up studies have largely involved altering the tokenization and self-attention modules to better adapt Transformers for addressing special challenges like non-stationarity, channel-wise dependency, and variable correlation in time series. However, we found that the expressive capability of sequence representation is a key factor influencing Transformer performance in time forecasting after investigating several representative methods, where there is an almost linear relationship between sequence representation entropy and mean square error, with more diverse representations performing better. In this paper, we propose a novel attention mechanism with Sequence Complementors and prove feasible from an information theory perspective, where these learnable sequences are able to provide complementary information beyond current input to feed attention. We further enhance the Sequence Complementors via a diversification loss that is theoretically covered. The empirical evaluation of both long-term and short-term forecasting has confirmed its superiority over the recent state-of-the-art methods.
Problem

Research questions and friction points this paper is trying to address.

Time Series Prediction
Transformer Models
Data Representation Optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Attention Mechanism
Transformer Model
Time Series Prediction
๐Ÿ”Ž Similar Papers
No similar papers found.