🤖 AI Summary
This study addresses the limitations of conventional time series forecasting methods, which often neglect latent dynamics and inadequately fuse multimodal covariates. To overcome these challenges, this work proposes a world model-based forecasting framework. Its core innovation lies in directly integrating multimodal covariates into latent space dynamics modeling for the first time, enabling the joint optimization of state formation and evolution. Furthermore, a two-stage training strategy is introduced: the model first learns covariate-conditioned latent dynamics, followed by training the observation decoder while keeping the learned dynamics frozen. Extensive experiments conducted across 21 real-world datasets demonstrate that the proposed approach significantly enhances time series forecasting performance in multimodal scenarios.
📝 Abstract
Time series forecasting is typically framed as learning a direct mapping from historical to future observations in the observation space. However, sequences of observations generally provide only a partial view of the dynamics of the underlying system, with future observations being shaped by latent dynamics. Recent latent-space forecasting methods thus achieve improved performance by predicting future observations from latent-space representations of historical observations rather than directly forecasting future observations in the observation space. Next, while future observations are also shaped by external factors, how to incorporate external, often multimodal, information into forecasting, so that it can shape latent-state formation and evolution directly, remains underexplored. We propose WorldTS, a world-modeling based forecasting framework that integrates multimodal covariates directly into the forecasting to further improve forecasting performance. Specifically, WorldTS employs a two-stage training strategy. First, it learns forecasting-relevant latent state dynamics conditioned on multimodal covariates, yielding encoded future states. Next, the learned state dynamics are frozen, and an observation decoder is trained to map the predicted future states back to future observations. Extensive experiments on 21 real-world datasets offer insight into WorldTS and its effectiveness.