Score
Design and build models, training procedures, and evaluation setups that forecast multiple future temporal chunks or horizons at once; this includes producing multi-horizon outputs, implementing mechanisms that form causal chains across prediction depths, and applying training strategies (e.g., next-forcing) and dense multi-scale intermediate supervision to fuse features from multiple layers and improve multi-step accuracy.
A core challenge in long-term time-series forecasting is the cumulative error growth with prediction horizon. This paper formally defines Long-Horizon Forecasting (LHF) as an error propagation problem and systematically surveys 35 years of research, emphasizing deep learning advances in trend/cyclic modeling, error suppression, and feature-space construction. We propose a unified framework integrating Fourier/wavelet transforms, bandpass filtering, State Space Models (SSMs), attention masking, sparse convolutions, and multi-scale decomposition. We identify xLSTM and Triformer as architectures capable of breaking the monotonic error-growth bottleneck. Furthermore, we release the first open-source LHF model library. Empirical evaluation on the ETTm2 dataset demonstrates that both models significantly mitigate MSE escalation: on the multivariate HUFL task, they achieve an average 18.7% reduction in forecasting error.
Multi-step time series forecasting (MSF) lacks a systematic theoretical foundation and evaluation framework for strategy selection and integration. This paper formally defines the MSF strategy space for the first time and proposes a learnable, composable, parameterized strategy framework that enables task-adaptive strategy search. We design a unified cross-strategy interface to decouple forecasting methods from strategy logic and support flexible integration. Extensive experiments—spanning 18 datasets, 5 base models, and 4 forecast horizons—yield 1,080 configurations; our framework achieves statistically significant improvements over state-of-the-art methods in over 84% of cases. To our knowledge, this work establishes the most comprehensive MSF strategy benchmark to date, providing a unified paradigm and empirical foundation for strategy modeling, evaluation, and automation in multi-step forecasting.
This work addresses the fundamental trade-off in long-term time series forecasting between parallel efficiency and temporal consistency: direct multi-step (DMS) approaches often lack sequential coherence, while iterative methods (IMS) suffer from error accumulation and slow inference. To overcome these limitations, we propose BTTF, a novel two-stage framework featuring lookahead enhancement and parallel self-refinement. BTTF leverages the model’s own predictions as augmentation signals and employs a lightweight ensemble to effectively refine a base forecaster. Notably, our approach requires no complex architectural modifications and is compatible with both linear and lightweight models. Extensive experiments demonstrate that BTTF achieves up to a 58% accuracy improvement across diverse long-horizon forecasting tasks, delivering consistent gains even when the base model is undertrained—thereby substantially transcending the inherent constraints of both DMS and IMS paradigms.
This work addresses the limitations of traditional time series forecasting methods, which treat the future as a static target and struggle to capture dynamically evolving non-stationary patterns. The study reframes forecasting as an adaptive multi-step scheduling problem—a novel perspective—and introduces a hierarchical controller that dynamically determines both the prediction scale and step size at each iteration. By integrating neural controlled differential equations, the approach enables hybrid continuous-discrete temporal evolution. This formulation substantially enhances modeling capacity for non-stationary dynamics, yielding average performance improvements of over 7.4% across multiple real-world and synthetic datasets. Moreover, the method achieves 2.6–5.3× faster inference than standard Transformer-based models and offers enhanced interpretability through explicit scheduling trajectories.
This work addresses the critical yet poorly understood role of long-horizon multi-turn planning in foundation model agents, which is hindered by the uncontrolled nature of internet-scale pretraining data. The authors construct a unified and controllable multi-turn environment to systematically investigate how the format, distribution, and quality of pretraining data influence planning capabilities. During post-training, they introduce GRPO, Online Policy Distillation (OPD), and Multi-teacher Online Policy Distillation (MOPD) to shape and integrate planning skills. Their findings reveal the essential role of explicit world models in enabling long-horizon generalization, demonstrate OPD’s superiority over GRPO under low-quality long-horizon data, and present the first successful fusion and transfer of cross-environment planning abilities via MOPD. Experiments further validate the efficacy of limited high-quality data and elucidate MOPD’s robust generalization, continual learning, and interference resilience across compatible, partially shared, or conflicting planning paradigms.
This work addresses the limited stability of conventional time series forecasting models, which rely solely on unidirectional inference from historical observations to future targets and neglect the structural information embedded in the unobservable trajectory beyond the prediction horizon. To overcome this limitation, we propose the KUP-BI paradigm, which for the first time treats the post-target continuation as a structured prior. Leveraging knowledge distillation, we construct a lightweight proxy of this continuation from a historical trajectory repository and introduce a feature-gated fusion module to enable bidirectional interaction between input features and the continuation proxy at the representation level. Notably, our approach requires no additional external data and integrates seamlessly with mainstream time series backbones, achieving significant state-of-the-art performance gains across six public benchmarks with minimal computational overhead.
为解决时间序列预测中模型随时间变化的有效性问题,本文提出TimEvolve方法,通过将每个实际结果转换为专家信任度、路径选择和干预强度的持久更新来系统地利用延迟反馈。
Pretrained time series foundation models often underperform on downstream tasks due to domain shift, task heterogeneity, scarce labeled data, and computational constraints. This work proposes the first systematic post-training framework, categorizing existing approaches along five dimensions based on their intervention points within the forecasting pipeline: parameter adaptation, context augmentation, model composition, output and uncertainty calibration, and compression with specialization. By delineating the design space and inherent limitations of each category, the framework offers a structured pathway to bridge the gap between pretraining and reliable deployment, thereby advancing the standardization and systematic development of time series post-training methodologies.
This work addresses the limitations of traditional world models that rely on recursive multi-step prediction, which often suffer from error accumulation and degraded long-horizon forecasting accuracy. The authors propose a novel end-to-end, non-recursive Direct Prediction World Model (DPWM) that compresses action sequences of arbitrary length into a single embedding and directly predicts the final state in a single forward pass. By avoiding recursive unrolling, DPWM mitigates error propagation and training instability while enabling gradient flow over extended horizons. Evaluated on continuous control and pixel-based benchmarks, the method consistently outperforms existing recursive models, with performance gains becoming more pronounced as the prediction horizon increases—demonstrating the critical role of terminal-state prediction as an effective objective for long-horizon tasks.
This work addresses the critical yet underexplored impact of deployment strategies on performance in multi-horizon volatility forecasting. The authors propose an adaptive deployment framework that dynamically selects rollout rules based on validation-set performance, automatically balancing prediction accuracy against inference cost across diverse model architectures—from linear models to PatchTST—and loss functions such as MSE and QLIKE. Experiments on volatility sequences of 20 stocks demonstrate that non-default rollout rules significantly outperform standard multi-output (MIMO) deployment. Notably, a small subset of rollout rules can closely approximate the performance of large ensembles at substantially lower computational overhead, with strategy effectiveness highly dependent on the choice of evaluation metric. This study underscores deployment strategy as a key source of adaptability in financial forecasting, challenging conventional fixed-deployment paradigms.
This study addresses the grid instability caused by the intermittent nature of solar irradiance in photovoltaic power generation. Existing approaches are often limited to single-step forecasting and tied to specific model architectures. To overcome these limitations, this work proposes an architecture-agnostic, multi-step joint prediction framework that simultaneously optimizes outputs across multiple future time steps, thereby enhancing the model’s capacity to capture long-term temporal dependencies. By integrating sequential sky images with historical power generation data, the framework is readily adaptable to diverse deep learning architectures. It achieves substantial improvements in prediction accuracy and robustness across all time horizons with negligible additional computational overhead, effectively strengthening grid resilience.