Pythia: Toward Foundation World Models for Multimodal Time Series

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of reusable representations and the difficulty of unified modeling across heterogeneous domains in multimodal time series forecasting. To this end, it proposes a world model based on the Joint Embedding Predictive Architecture (JEPA). Methodologically, the approach employs stop-gradient numerical references to guide context rectification, thereby decoupling world model pretraining from the probabilistic decoder. Furthermore, it integrates textual and auxiliary information to achieve multimodal contextual complementarity, optimizing the learning of temporal dynamics. Experimental results demonstrate that the proposed method reduces prediction error by 5%–6.26% compared to the strongest baseline on the MUSE benchmark, validating the effectiveness of shared pretraining and multimodal fusion for time series forecasting.
📝 Abstract
Time-series foundation models offer a unified approach to forecasting across heterogeneous domains. Textual context and auxiliary observations provide complementary information about temporal dynamics, yet reusable multimodal predictive representations remain underexplored. We introduce Pythia, a foundation world model that learns context-conditioned latent dynamics across datasets through a joint-embedding predictive architecture. A stop-gradient numerical reference guides contextual corrections to predicted future states. A separate probabilistic decoder then adapts to the frozen predictive representation and observed history, decoupling world-model pretraining from observation-space forecasting. On MUSE, Pythia-Tiny's normalized mean absolute scaled error (MASE) and weighted sum quantile loss (WSQL) are 0.6879 and 0.4269, reducing errors by 6.26% and 5.00% relative to the strongest model evaluated in the published MUSE leaderboard. Through a series of controlled experiments, we investigate how to design a time-series world model through shared pretraining and how joint-embedding predictive learning can incorporate multimodal information. The results support separating predictive representation learning from probabilistic readout and show complementary contributions from entity descriptions, events, and covariates.
Problem

Research questions and friction points this paper is trying to address.

Time-series foundation models
Multimodal time series
World models
Predictive representations
Joint-embedding predictive architecture
Innovation

Methods, ideas, or system contributions that make the work stand out.

Foundation World Model
Joint-Embedding Predictive Architecture
Multimodal Time Series
Stop-Gradient Mechanism
Probabilistic Decoder
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xilin Dai
Ant International
H
Hongzhou Chen
Ant International
Y
Yifan Hu
Ant International
Yiding Liu
Yiding Liu
TikTok
Z
Zewei Dong
Ant International
Jiang-Ming Yang
Jiang-Ming Yang
Ant Financial
PaymentsInformation RetrieverConsistency Maintenance