TrajSceneLLM: A Multimodal Perspective on Semantic GPS Trajectory Analysis

📅 2025-06-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the limitations of shallow semantic representation and insufficient spatial context integration in GPS trajectory modeling, this paper proposes the first multimodal trajectory representation framework that jointly leverages visual map images and large language model (LLM)-generated temporal motion text. The method eliminates hand-crafted feature engineering: a CNN encodes geospatial layout from map images, while an LLM captures high-level motion semantics from trajectory sequences; their embeddings are fused to model deep spatiotemporal dependencies. A subsequent embedding concatenation layer and MLP classifier enable end-to-end learning. Evaluated on transportation mode identification, the framework achieves significant performance gains over state-of-the-art baselines, empirically validating the effectiveness of multimodal spatiotemporal semantic embedding. The source code and benchmark dataset are publicly released.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Multimodal LearningNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
GPS trajectory data reveals valuable patterns of human mobility and urban dynamics, supporting a variety of spatial applications. However, traditional methods often struggle to extract deep semantic representations and incorporate contextual map information. We propose TrajSceneLLM, a multimodal perspective for enhancing semantic understanding of GPS trajectories. The framework integrates visualized map images (encoding spatial context) and textual descriptions generated through LLM reasoning (capturing temporal sequences and movement dynamics). Separate embeddings are generated for each modality and then concatenated to produce trajectory scene embeddings with rich semantic content which are further paired with a simple MLP classifier. We validate the proposed framework on Travel Mode Identification (TMI), a critical task for analyzing travel choices and understanding mobility behavior. Our experiments show that these embeddings achieve significant performance improvement, highlighting the advantage of our LLM-driven method in capturing deep spatio-temporal dependencies and reducing reliance on handcrafted features. This semantic enhancement promises significant potential for diverse downstream applications and future research in geospatial artificial intelligence. The source code and dataset are publicly available at: https://github.com/februarysea/TrajSceneLLM.
Problem

Research questions and friction points this paper is trying to address.

Extracting deep semantic representations from GPS trajectories
Incorporating contextual map information into trajectory analysis
Improving performance in travel mode identification tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal integration of GPS and map images
LLM-generated textual descriptions for trajectories
Simple MLP classifier with rich semantic embeddings
🔎 Similar Papers
No similar papers found.
C
Chunhou Ji
The Hong Kong University of Science and Technology (Guangzhou)
Q
Qiumeng Li
The Hong Kong University of Science and Technology (Guangzhou)