Score
Designs, implements, and evaluates methods that convert trajectories—ordered spatiotemporal sequences of positions, states, or features—into compact, informative representations or encodings (embeddings) that preserve temporal and spatial structure. Builds algorithms to infer latent trajectory structure, parse and segment trajectories into meaningful sub-trajectories, apply regularization, and produce summaries for downstream tasks such as prediction, similarity search, compression, or visualization.
This work addresses the problem of learning task-agnostic trajectory representations from state-action sequences without reward annotations. Methodologically, inspired by CLIP and BERT, we propose the first multi-capability-compatible trajectory embedding framework that jointly leverages contrastive learning and self-supervised sequence modeling, augmented with latent-geometric constraints to enforce behavioral controllability and additive structure in the embedding space. Crucially, the framework requires no reward signals, yet learns general-purpose trajectory representations supporting diverse downstream tasks—including imitation learning, behavior classification, clustering, and trajectory regression. Extensive evaluation across autonomous driving, robotics, and healthcare benchmarks demonstrates significant improvements over prior methods, validating strong cross-task and cross-domain generalization. The implementation is publicly available.
Retrieving short trajectories with dual similarity in semantics and direction remains challenging due to the insensitivity of conventional metrics (e.g., FFT) to directional characteristics. Method: This paper proposes a lightweight contrastive learning embedding framework leveraging a Transformer encoder jointly optimized with triplet loss, and introduces— for the first time—a Cosine-based contrastive objective incorporating directional intent modeling. The learned embeddings achieve high discriminability and interpretability while supporting real-time inference, with dimensionality as low as 4D. Results: Evaluated on Argoverse 2, the method significantly improves minADE and minFDE. Even at 4D, it maintains superior retrieval performance while drastically reducing computational overhead, enabling real-time motion prediction and deployment in autonomous navigation systems.
This work addresses the problem of low-dimensional embedding for temporal network trajectories. We propose a scalar embedding paradigm centered on preserving inter-snapshot relative graph distances—departing from conventional approaches that rely solely on static snapshot topology. Our method models dynamic networks as trajectories in a graph metric space and achieves distance-preserving mapping into a one-dimensional Euclidean space via multidimensional scaling (MDS) and principal component analysis (PCA) applied to graph distance matrices. We employ robust graph distance measures—including Gromov–Wasserstein distance and delta-convergence—to ensure stability under structural perturbations. Experiments on both synthetic and real-world datasets demonstrate that the resulting 1D scalar sequence accurately captures complex dynamical patterns such as periodicity, abrupt transitions, and decay. This significantly enhances feasibility, interpretability, and computational efficiency of temporal network analysis, and—crucially—constitutes the first effective representation of dynamic network structural evolution in a 1D embedding.
To address the limitations of shallow semantic representation and insufficient spatial context integration in GPS trajectory modeling, this paper proposes the first multimodal trajectory representation framework that jointly leverages visual map images and large language model (LLM)-generated temporal motion text. The method eliminates hand-crafted feature engineering: a CNN encodes geospatial layout from map images, while an LLM captures high-level motion semantics from trajectory sequences; their embeddings are fused to model deep spatiotemporal dependencies. A subsequent embedding concatenation layer and MLP classifier enable end-to-end learning. Evaluated on transportation mode identification, the framework achieves significant performance gains over state-of-the-art baselines, empirically validating the effectiveness of multimodal spatiotemporal semantic embedding. The source code and benchmark dataset are publicly released.
This work addresses the challenges posed by raw GPS trajectories—characterized by continuity, high noise, and irregular sampling—which hinder traditional spatial tokenization methods from simultaneously achieving fine-grained representation and discriminative pattern modeling. The authors propose TrajTok, a novel framework that learns multi-resolution hexagonal grids from trajectory point distributions to enable adaptive spatial tokenization. Coupled with a factorized Transformer encoder and masked token pretraining, TrajTok produces transferable trajectory representations. Innovatively integrating data-driven spatial partitioning, spatiotemporal rotary position encoding (ST-RoPE), and cross-modal attention mechanisms, the method achieves state-of-the-art performance across diverse downstream tasks—including trajectory similarity search, classification, arrival time estimation, and full-trip duration regression—on the Porto dataset, using only a frozen encoder with lightweight adapters.
Understanding whether language-model embeddings encode structured real-world information is important for both representation analysis and information retrieval. We study this question for temporal and geographic signals using a simple projection-based method that operates directly on output embeddings. Given a small set of seed examples, the method defines an axis in embedding space and ranks texts or entities by their projection onto that axis. Our approach is fully black-box and model-agnostic: it requires only embeddings, without access to model weights, internal activations, auxiliary probes, or additional training. This makes it applicable to modern embedding models available only through APIs and provides a lightweight way to analyze whether temporal and spatial dimensions are present in their representation spaces. We apply the method to temporal and geographic datasets and find that embedding projections recover meaningful chronological and spatial structure. These results provide evidence that output embeddings encode signals relevant to time and space, while also offering a practical tool for interpretability and for downstream temporal and geographic information retrieval tasks, such as temporal ordering, geographic ranking, and tagging.
This work addresses time series characterized by irreversible state evolution—such as equipment degradation, task completion, or neural dynamics—and proposes a novel “latent compass” representation. By leveraging self-supervised contrastive learning, the method constructs a structured latent space in which each time series is mapped onto a manifold trajectory between two orthogonal prototype vectors. State progression and operational mode are disentangled via polar coordinates (θ, r), enabling transparent and interpretable modeling without requiring labeled data. Evaluated on industrial degradation, robotic tasks, and neural activity datasets, the approach achieves performance on par with or superior to black-box deep models in endpoint prediction, multi-step forecasting, and phase separation tasks—even when paired with simple linear regression—while substantially enhancing model interpretability and computational efficiency.
Existing learning-based approaches for trajectory similarity lack theoretical guarantees, exhibit weak generalization, and incur high training costs. This work proposes LB-TrajRep, the first unified trajectory lower-bound representation framework that dispenses with neural embeddings. Leveraging a point-pivot mechanism, LB-TrajRep generates single-vector representations that yield interpretable lower bounds for multiple classical distance measures—including Dynamic Time Warping (DTW), Hausdorff, and discrete Fréchet distance—and seamlessly integrates into standard vector retrieval pipelines. Two data-driven pivot selection strategies are introduced to optimize either bound tightness or hard-sample ranking quality. Experiments on real-world trajectory datasets demonstrate that LB-TrajRep significantly outperforms existing neural embedding methods, achieving 20%–60% gains in Top-k ranking accuracy for Hausdorff and discrete Fréchet distances and 15%–40% improvements for DTW.
研究通过分析V-JEPA 2和VideoMAE-v2模型,探讨了视频基础模型中时空表示的编码内容、出现位置及几何组织方式,并使用轻量级探针来发现三种时间属性。
This work proposes TrajLoom, a method for predicting dense, long-horizon future point trajectories and their visibility from observed videos to support video understanding and controllable generation. To mitigate positional bias, the approach introduces Grid-Anchor Offset Encoding and designs TrajLoom-VAE, which learns a structured latent trajectory space through masked reconstruction and spatiotemporal consistency regularization. Long-range generation stability is achieved via TrajLoom-Flow, leveraging flow matching, boundary-aware prompts, and on-policy K-step fine-tuning. The study also establishes TrajLoomBench, a unified benchmark for evaluation. Experiments demonstrate that TrajLoom extends the prediction horizon from 24 to 81 frames, significantly improving trajectory realism and temporal stability across multiple datasets, with generated trajectories directly applicable to downstream video synthesis and editing tasks.