🤖 AI Summary
Existing time series forecasting models rely on paradigm-specific designs, hindering the unified processing of numerical and narrative streaming data. This work proposes MIDAPN, a unified spatiotemporal forecasting backbone that decouples media identities via a multimedia identity-aware graph and balances multiscale temporal features through spectral prism convolutions. By integrating static, dynamic, and latent commonality modeling with context-aware identity modulation and search-guided adaptive architecture configuration, the framework achieves universal graph adaptation for cross-modal signals and automated temporal learning. Extensive comparative experiments across 25 datasets against 30 mainstream baselines demonstrate that MIDAPN delivers superior predictive performance while maintaining robust compatibility as a general-purpose backbone network.
📝 Abstract
Media-bridged time series forecasting is expanding to encompass traditional "multivariate" and emerging "multimodal" (e.g., through textual assistance). Existing Time Series Forecasting (TSF) models still rely on paradigm-specific relation, fusion, and temporal modules, hindering a common forecasting backbone across numerical and pre-aligned narrative-flow settings. To explore this, we propose the Multimedia Identity-Aware Prism Network (MIDAPN), a unified spatiotemporal forecasting backbone based on media-general graph adaptation and automatic temporal learning: (1) Following media pre-alignment, our Multimedia Identity-Aware Graph (MIDAG) revisits identity through static essence, dynamic behavior, and latent commonality, inducing affinities that extend variable-specific dependencies across media. Contextual Identity Modulation (CIM) further refines discriminative aggregation. (2) We develop Spectral Prism Convolution (SPConv) to automatically perform hierarchical temporal analysis, balancing coarse trends and fine-grained details. Meanwhile, its Adaptive Search Guidance configures a scale-efficient architecture for temporal-dimension reconstruction. These decoupled yet synergistic components jointly address media identity disentanglement and temporal-scale mismatch. Comprehensive evaluations involving 16 SOTA TSF models across 13 "multivariate" and 12 "multimodal" datasets, alongside targeted long-context comparisons against 14 time series foundation models and fused pretrained language models, demonstrate MIDAPN's consistent superiority and broad shared backbone compatibility. The code is available at \href{https://github.com/leijieruilq/MIDAPN/tree/main}{https://github.com/MIDAPN}.