Aurora-X: Built for Extreme Time Series Forecasting

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited training potential and architectural inflexibility of general-purpose time series foundation models by constructing a billion-parameter foundation model. Methodologically, it proposes a pattern-guided Mixture-of-Experts mechanism and an implicit quantile network head to enable sparsely activated routing and arbitrary quantile probabilistic forecasting. By integrating channel-independent pretraining, progressive curriculum learning, and variable-resolution post-training, the model supports flexible inference across variates, covariates, and multiple resolutions, while incorporating test-time scaling and parallel decoding to enhance efficiency. The proposed model achieves state-of-the-art performance on benchmarks such as GIFT-Eval, comprehensively outperforming existing pretrained and task-specific supervised models.
📝 Abstract
Time series foundation models (TSFMs) enable cross-domain forecasting, but their development as general-purpose forecasters remains constrained by underexplored training potential and limited architectural versatility. To address these challenges, we introduce Aurora-X, a billion-scale TSFM with a progressive curriculum and a unified architecture. We first use channel-independent pretraining to learn temporal patterns, then introduce cross-variable dependencies, varied context and horizon lengths, and future covariates if available during midtraining. Variable-resolution post-training further enables an adjustable temporal span per token at inference. With fixed model weights, this supports longer histories under a fixed token budget or fewer tokens for the same history, enabling test-time scaling. With a versatile architecture, Aurora-X supports cross-variable modeling, covariate conditioning, and parallel decoding of future patches for probabilistic forecasting. These are supported by a novel pattern-guided mixture-of-experts that expands model capacity through sparse activation and uses shallow patch similarities to constrain deep-layer routing, guiding expert specialization across heterogeneous time series. Furthermore, we propose an implicit quantile network head that predicts arbitrary quantiles to characterize predictive distributions, enhancing probabilistic forecasting flexibility. Comprehensive experiments on GIFT-Eval, TIME, FEV-Bench, TFB, and DAG-Bench demonstrate state-of-the-art forecasting performance against pretrained TSFMs and task-specific supervised models.
Problem

Research questions and friction points this paper is trying to address.

Time Series Foundation Models
Cross-domain Forecasting
Architectural Versatility
Probabilistic Forecasting
Innovation

Methods, ideas, or system contributions that make the work stand out.

Time Series Foundation Model
Progressive Curriculum Training
Pattern-guided Mixture-of-Experts
Implicit Quantile Network
Variable-resolution Post-training
🔎 Similar Papers
No similar papers found.