Score
Designs and implements self-supervised learning methods that produce compact embeddings of trajectories (sequences of states or timepoints) from unlabeled data. Builds pretext objectives and training pipelines—such as temporal-consistency losses, contrastive or predictive tasks, and complementary regularizers—to learn representations that improve downstream classification or analysis of trajectory behavior.
This work addresses the problem of learning task-agnostic trajectory representations from state-action sequences without reward annotations. Methodologically, inspired by CLIP and BERT, we propose the first multi-capability-compatible trajectory embedding framework that jointly leverages contrastive learning and self-supervised sequence modeling, augmented with latent-geometric constraints to enforce behavioral controllability and additive structure in the embedding space. Crucially, the framework requires no reward signals, yet learns general-purpose trajectory representations supporting diverse downstream tasks—including imitation learning, behavior classification, clustering, and trajectory regression. Extensive evaluation across autonomous driving, robotics, and healthcare benchmarks demonstrates significant improvements over prior methods, validating strong cross-task and cross-domain generalization. The implementation is publicly available.
This work addresses an inherent tension in self-supervised spatiotemporal forecasting: recurrent models process frames sequentially but ignore local redundancy, whereas non-recurrent models stack frame sequences yet lose temporal structure. To resolve this, we propose a dual-scale unified modeling framework—the first to jointly integrate recurrent temporal modeling (via hidden-state recurrence) and non-recurrent temporal modeling (via cross-frame global attention)—to simultaneously capture fine-grained frame-level dynamics and coarse-grained sequence-level dependencies. We further introduce a hierarchical spatiotemporal masked prediction objective and contrastive regularization. Evaluated on multiple standard benchmarks, our method achieves significant improvements over state-of-the-art approaches, reducing average prediction error by 12.7%, while also enhancing generalization and computational efficiency.
In pedestrian trajectory prediction, supervised learning suffers from long-tailed data distributions and struggles to model anomalous behaviors such as abrupt stops or sharp turns. To address this, we propose the first self-supervised framework that explicitly and jointly models position, velocity, and acceleration. Our method introduces a hierarchical velocity/acceleration feature injection architecture, enforces physical kinematic consistency via a novel self-supervised mechanism, and integrates a pseudo-label generation strategy to enable cooperative prediction and dynamic coupling constraints among the three motion variables. Crucially, the framework requires no ground-truth velocity or acceleration annotations—only raw trajectory coordinates are needed for training. Evaluated on ETH-UCY and Stanford Drone datasets, our approach achieves state-of-the-art performance, demonstrating significant improvements in both prediction accuracy and robustness for anomalous motions.
This paper addresses the challenging problem of identifying latent dynamical systems from nonlinear observations. We propose Dynamics Contrastive Learning (DCL), the first framework to theoretically establish that self-supervised contrastive learning enables identifiable recovery of latent dynamics. Methodologically, DCL operates in a fully unsupervised manner—requiring neither labels nor prior assumptions about dynamical structure—and disentangles linear, switched-linear, and nonlinear (including chaotic) latent dynamics by constructing dynamics-consistent positive and negative sample pairs directly from nonlinear observational data. Key contributions include: (1) establishing a rigorous theoretical connection between self-supervised learning and causal generative factor disentanglement; (2) providing the first identifiability guarantee for self-supervised learning–driven system identification; and (3) demonstrating high-fidelity reconstruction across diverse dynamical regimes on both synthetic and benchmark dynamical datasets.
This work addresses the challenge of modeling multimodal behaviors in context-free 2D trajectory prediction by proposing an unsupervised, self-conditioned generative adversarial network (GAN) that requires no external scene information. The method implicitly captures diverse motion patterns through the discriminator’s feature space and incorporates three tailored training strategies to enhance both diversity and accuracy of predictions. As the first study to apply self-conditioned GANs to context-free trajectory forecasting, the model consistently outperforms existing context-free approaches on both human motion and road-agent datasets, demonstrating particularly strong performance on sparsely labeled categories and achieving state-of-the-art results in human motion prediction.
This work addresses a critical limitation in self-supervised dynamic representation learning, where existing contrastive predictive objectives often misinterpret slowly varying noise within trajectories as genuine dynamical signals, leading to noise-dominated representations and degraded downstream performance. The authors identify this issue as stemming from an inherent inductive bias flaw in standard contrastive objectives and propose a general corrective principle: sampling negative examples from within the same trajectory to eliminate predictive shortcuts introduced by slow-varying noise, thereby compelling the encoder to focus on the true dynamical variables governing system evolution. Experiments based on frameworks such as JEPA and DySIB on synthetic moving-point and rigid-pendulum video datasets demonstrate that the proposed approach effectively disentangles slow noise from authentic dynamics, yields representations whose quality improves with trajectory length, and significantly enhances downstream task performance under strong noise conditions.
This work addresses the limitations of handcrafted augmentation strategies in time series contrastive learning, which often introduce spurious correlations and suffer from poor generalization. To overcome these issues, the authors propose a novel paradigm that explicitly encodes temporal shift invariance to construct deterministic views, thereby replacing conventional domain-knowledge-dependent augmentations. This approach leverages temporal shift invariance alone to generate effective positive and negative sample pairs, significantly reducing reliance on manual intervention. Evaluated across six real-world benchmarks and the UCR/UEA archive, the method achieves state-of-the-art performance while substantially accelerating training. Furthermore, the study systematically investigates the impact of batch size and the number of negative samples on model effectiveness, offering valuable insights into the design of contrastive learning frameworks for time series data.
This work addresses the significant performance degradation of autonomous driving trajectory prediction models under test-time distribution shifts, which often leads to erroneous predictions in unfamiliar scenarios. The authors propose a self-supervised post-processing method that requires no modification to the original model: a decoder is trained to predict the latter half of a trajectory from its first half, and the L2 norm of the gradient of the prediction loss with respect to the decoder’s final-layer parameters serves as a distribution shift detection score. This approach efficiently identifies out-of-distribution inputs and enables early collision warnings. Experiments demonstrate that the method substantially outperforms existing approaches on the Shifts and Argoverse datasets and has been successfully integrated into a Deep Q-Network motion planner within the Highway simulation environment, achieving reliable collision risk detection.
This work addresses time series characterized by irreversible state evolution—such as equipment degradation, task completion, or neural dynamics—and proposes a novel “latent compass” representation. By leveraging self-supervised contrastive learning, the method constructs a structured latent space in which each time series is mapped onto a manifold trajectory between two orthogonal prototype vectors. State progression and operational mode are disentangled via polar coordinates (θ, r), enabling transparent and interpretable modeling without requiring labeled data. Evaluated on industrial degradation, robotic tasks, and neural activity datasets, the approach achieves performance on par with or superior to black-box deep models in endpoint prediction, multi-step forecasting, and phase separation tasks—even when paired with simple linear regression—while substantially enhancing model interpretability and computational efficiency.
This work addresses the challenges of high computational cost, low sample efficiency, and poor generalization in robotic trajectory planning within high-dimensional, cluttered environments. To overcome these limitations, we propose a neuro-inspired self-supervised learning framework that jointly trains forward and inverse dynamics models, leveraging intrinsic supervisory signals instead of expert demonstrations or extensive environmental exploration. A novel training strategy is introduced to mitigate over-reliance on learned signals, enhancing model robustness. Experimental results demonstrate that the proposed approach significantly improves planning efficiency, performance, and robustness in complex obstacle-rich scenarios, enabling effective and generalizable trajectory generation.