T-JEPA: A Temporal Joint-Embedding Predictive Architecture for Learning Better Remote Sensing Representations

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of leveraging sparse, irregularly sampled temporal observations in Earth observation as supervision signals for remote sensing foundation models. We propose T-JEPA, a joint-embedding predictive architecture that learns latent transitions conditioned on time intervals, predicting complete target latent fields from masked source representations and true temporal spans to organize discrete observations into structured latent trajectories. The method introduces asymmetric metadata injection to mitigate shortcut learning and adopts direct cross-scale supervision instead of recursive unrolling strategies. Experiments demonstrate that T-JEPA achieves state-of-the-art transfer performance on both static and temporal tasks, yielding predictable transitions with explicit interval dependence and strong semantic discriminability.
📝 Abstract
Earth observation (EO) data provide rich temporal supervision, yet existing remote sensing foundation models mainly exploit sequential observations through imposing predefined pairwise relations or aggregating holistic reconstruction context. We seek to further exploit the sparse and nonuniform temporal sampling inherent in EO sequences as supervisory signals. To this end, we propose T-JEPA, a temporal joint-embedding predictive architecture that learns time-gap-conditioned latent transitions. A shared single-frame encoder processes each observation, while a temporal predictor estimates the complete target latent field from a masked source latent representation and the actual elapsed time. Across multiple temporal intervals, these predictive constraints organize observed states into structured latent trajectories. Asymmetric metadata injection mitigates shortcut learning, and direct supervision across multiple temporal scales proves more effective than recursively rolling out intermediate states. In parallel, masked pixel reconstruction provides complementary supervision for preserving spatial details. Under matched pre-training data and throughput, T-JEPA achieves leading transfer performance on both static and temporal tasks. Analyses further reveal that T-JEPA learns representations with time-gap-dependent transition predictability and coherent latent dynamics, while maintaining strong cross-period consistency, representation diversity, and semantic discriminability.
Problem

Research questions and friction points this paper is trying to address.

Remote Sensing
Foundation Models
Temporal Representation Learning
Earth Observation
Self-Supervised Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Temporal Joint-Embedding Predictive Architecture
Time-gap-conditioned Latent Transitions
Asymmetric Metadata Injection
Multi-scale Direct Supervision
Masked Pixel Reconstruction
🔎 Similar Papers
No similar papers found.