Score
Designs and evaluates encoders, loss functions, and training pipelines that produce joint spatial–temporal (and often cross‑modal) embeddings by aligning or contrasting representations across space and time; methods include contrastive objectives, embedding alignment, and cross‑modal reconstruction to fuse spatial context and temporal dynamics into compact feature vectors. Practically this skill builds models and alignment procedures whose outputs can be analyzed or used as inputs for downstream tasks (prediction, classification, conditioning, or model guidance) and includes measuring representation quality and alignment across modalities and time.
To address the task-specificity and poor generalizability of existing spatio-temporal deep learning models, this paper systematically surveys the full lifecycle of Spatio-Temporal Foundation Models (STFMs) and introduces the first structured pipeline framework. We propose a novel pipeline-oriented survey paradigm, establish a new taxonomy of data attributes tailored to spatio-temporal characteristics, and explicitly distinguish two core phases: “raw pretraining” and “downstream adaptation.” Within this framework, we unify the design principles for spatio-temporal embeddings, model architectures, pretraining objectives, and adaptation strategies. The framework supports data-driven modeling, multi-source dependency characterization, and multi-objective joint training—thereby significantly enhancing model reusability and development efficiency. It provides a reproducible methodological guide for STFM research and outlines concrete directions for future work.
This paper addresses the semantic alignment challenge in non-visual time-series-to-high-fidelity-image generation. We propose TimeArtist, a novel framework that introduces two key innovations: (1) a “warm-up–alignment” two-stage training paradigm—first performing self-supervised pretraining via dual autoencoders with a shared vector quantizer, then freezing encoders and introducing a representation-layer projection to achieve time-vision cross-modal semantic alignment; and (2) a transferable spatial prior unified architecture that maps temporal dynamics into controllable visual styles. Experiments demonstrate state-of-the-art performance on image quality metrics (e.g., FID, LPIPS) and achieve SOTA results on downstream zero-shot time-series forecasting tasks. These findings validate the effectiveness of both cross-modal semantic alignment and generative representation transfer.
This study investigates whether time series share a unified latent representational structure with visual and linguistic modalities and explores the limits of multimodal alignment. By applying contrastive learning to post-hoc align frozen time-series, vision, and language encoders, the work systematically analyzes the geometry of their representations, scaling behaviors, and dependence on information density. It reveals, for the first time, the role of time series in multimodal alignment: larger model scales improve overall alignment performance; time series align more readily with vision than with text; images serve as a mediating bridge across modalities; and both textual and visual modalities exhibit saturation thresholds beyond which increased information density yields diminishing returns.
This study systematically investigates video-text cross-modal representation alignment, focusing on modern encoders’ spatiotemporal modeling capabilities and their relationship with downstream performance. Method: We propose parameterized test-time scaling laws to quantitatively link semantic alignment degree with video understanding ability; design a novel temporal reasoning benchmark to overcome limitations of conventional zero-shot classification evaluation; and integrate multi-frame video encoding, text-set alignment, and regression-based modeling to jointly learn static and dynamic representations. Results: Experiments demonstrate that strong text alignment significantly enhances general-purpose video representation quality, and alignment metrics reliably predict model performance across diverse video understanding tasks. Our work advances the understanding of multimodal model internals and provides an interpretable pathway for alignment optimization.
Multimodal contrastive learning often captures only redundant, shared information across modalities, failing to model modality-unique and synergistic interactions. To address this, we propose CoMM, a framework that abandons explicit cross-modal feature alignment and instead maximizes mutual information among augmented representations within a unified multimodal embedding space—enabling end-to-end co-modeling. For the first time, we rigorously disentangle multimodal information into redundant, unique, and synergistic components from an information-theoretic perspective, and theoretically prove that mutual information maximization inherently balances these three components. The method is fully differentiable and requires no paired multimodal supervision. Controlled ablation studies validate the accuracy of our information disentanglement, and CoMM achieves state-of-the-art performance on seven real-world multimodal benchmarks.
To address the challenges of complex feature fusion and poor scalability in unsupervised multimodal time-series representation learning, this paper proposes a spectral-graph-theoretic framework for implicit cross-modal alignment. Methodologically, it abandons explicit multimodal feature fusion and instead adopts a single-encoder, multi-view architecture: diverse time-series views—such as frequency-domain, image-based, and symbolic representations—are generated via modality-agnostic temporal transformations, and spectral-graph-guided contrastive loss enables unsupervised cross-modal alignment. This lightweight design implicitly captures latent inter-modal dependencies, strengthening inductive bias and generalization capacity. Extensive experiments across multiple domains demonstrate state-of-the-art performance on downstream tasks—including classification, forecasting, and anomaly detection—surpassing existing unsupervised methods by an average accuracy gain of 3.2%–7.8%.
This work addresses the lack of a systematic understanding of when cross-modal alignment (CA) and cross-modal prediction (CP) are effective in multimodal learning—a gap that often leads to suboptimal performance or even degradation relative to unimodal baselines. The authors propose a unified linear analytical framework under a structured signal–noise model with correlated interference, revealing complementary failure mechanisms of CA and CP. They introduce the first multimodal “phase diagram,” which delineates four distinct regimes: both methods succeed, only alignment works, only prediction works, or neither is effective. Leveraging separation ratio analysis, a unidirectional whitening mechanism, and a few-shot label-guided localization algorithm, this phase diagram enables practical guidance for method selection on real-world data. Experiments across synthetic, stereo vision, image–text, and astrophysical datasets validate its efficacy in identifying harmful multimodal configurations, offering a diagnostic tool for practitioners prior to model deployment.
This paper addresses the cross-modal semantic gap in multimodal understanding through a unified alignment–translation–fusion–transfer framework. Methodologically: (1) a spatial reasoning BERT is introduced to map spatial language to 2D layouts; (2) a medical term spatial co-occurrence loss is designed to ground textual descriptions in 3D anatomical locations; (3) a structured text-to-knowledge graph fact linking benchmark with interpretability is established; and (4) a multi-stream feature fusion mechanism coupled with cross-modal knowledge distillation enables lightweight RGB-based action recognition. Key contributions include: the first spatial semantic alignment model, joint anatomical-spatial representation learning, a standardized, interpretable knowledge graph linking benchmark, and a novel unimodal distillation paradigm that achieves near-fused performance without multimodal inputs. Experiments demonstrate significant improvements across all tasks: the RGB-only model attains accuracy comparable to multimodal baselines while reducing computational overhead by over 60%.
This work addresses the challenges of heterogeneity neglect and information asymmetry in multimodal sentiment analysis caused by conflated spatiotemporal modeling. It proposes a novel approach that explicitly decouples each modality into temporal dynamics and spatial structure representations. The framework employs a dual time–space encoder, factor-consistent cross-modal alignment, factor-specific supervision, and decorrelation regularization, followed by a gated re-coupling module for effective fusion. By introducing factor-level alignment and a leakage-prevention mechanism, the method enhances model interpretability while achieving state-of-the-art performance across multiple benchmark datasets. Ablation studies further confirm the effectiveness and necessity of each component in the proposed architecture.
This work addresses two key challenges in video grounding—tight spatiotemporal alignment coupling and visual token redundancy—by proposing the Bridge-STG framework, which decouples temporal and spatial localization tasks to enable heterogeneous subtask optimization while preserving semantic consistency. The core innovations include a Spatio-Temporal Semantic Bridging (STSB) mechanism to bridge the semantic gap introduced by decoupling, and a Query-Guided Spatial Localization (QGSL) module to eliminate redundancy across both domains. Integrated with explicit temporal alignment, multi-layer interactive queries, positive-negative frame sampling, and end-to-end multitask training, the method achieves state-of-the-art performance among multimodal large language models on VidSTG, improving m_vIoU from 26.4 to 34.3, and demonstrates strong cross-task transferability.