Score
Design and build models, algorithms, and datasets for spatio-temporal trajectories, including prediction, forecasting, generation, synthesis, sampling, and scoring of agent or object paths. Also develop pipelines and tools for trajectory annotation and logging, data collection and processing, mining and replay, planning and tracing, and evaluation metrics to validate and improve trajectory-model behavior.
Addressing the challenge of simultaneously ensuring diversity, plausibility, and interpretability in vision-driven multi-trajectory prediction (MTP), this paper presents a systematic survey of the field. We propose the first unified taxonomy for MTP, comprehensively categorizing model paradigms—including variational autoencoders (VAEs), generative adversarial networks (GANs), diffusion models, graph neural networks (GNNs), and attention mechanisms—alongside mainstream open-source datasets and evaluation metrics. By analyzing the interplay between uncertainty modeling and interpretability, we identify key technical bottlenecks and distill five emerging research directions. This work establishes a principled theoretical framework and practical guidelines for developing safe, reliable autonomous navigation systems. (128 words)
To address the lack of standardized datasets, inconsistent preprocessing protocols, and non-uniform evaluation metrics in UAV trajectory prediction research, this paper proposes the first comprehensive standardization framework for the field. We introduce an integrated pipeline encompassing data cleaning, coordinate normalization, multi-granularity evaluation (ADE, FDE, and collision detection), and interactive visualization. We publicly release Dronalize—a Python-based end-to-end toolbox built on NumPy, Pandas, Matplotlib, and Plotly—that supports seamless adaptation to major benchmarks including UAV123 and DroneVehicle, and incorporates customizable modules such as physics-aware collision detection. Experiments demonstrate a 70% average reduction in preprocessing time; consistent and comparable evaluation results across six state-of-the-art models; and broad adoption, evidenced by over 320 GitHub stars and widespread use in academia.
A systematic review and methodological synthesis of Trajectory Foundation Models (TFMs)—a pivotal subclass of Spatio-Temporal Foundation Models (STFMs)—is currently lacking. This paper introduces the first unified taxonomy for TFMs, integrating spatio-temporal representation learning, pretraining-finetuning paradigms, and multi-task transfer learning to systematically analyze model architectures, training strategies, and task adaptation mechanisms. Our contributions are threefold: (1) We propose the first comprehensive methodological taxonomy for TFMs, explicitly identifying core challenges—including trajectory data sparsity, spatio-temporal heterogeneity, and generalization robustness; (2) We delineate a principled technical pathway toward transferable, trustworthy, and general-purpose spatio-temporal intelligence; and (3) We distill high-level research directions in spatio-temporal knowledge modeling, with direct implications for urban computing, traffic forecasting, and mobile intelligence applications.
A systematic methodology and reproducible practice for building mobile trajectory foundation models (TrajFMs) from scratch remains lacking. Method: We propose a lightweight TrajFM construction paradigm based on the GPT-2 architecture, introducing time-series patching to trajectory modeling for the first time, and designing spatiotemporal-aware data encoding and positional embedding strategies. We release a fully modular, open-source implementation. Within a unified evaluation framework, we comparatively analyze TrajFM, TrajGPT, and other state-of-the-art methods, clarifying fundamental differences in modeling assumptions, input representations, and training objectives. Contribution/Results: This work bridges critical pedagogical and engineering gaps in trajectory AI, substantially enhancing transparency, auditability, and reproducibility in TrajFM development. It provides a standardized technical reference for the SIGSPATIAL and broader spatial AI communities.
High-quality multi-step GUI interaction trajectories for training GUI agents are scarce and prohibitively expensive to annotate manually. Method: This paper proposes a web-tutorial-based automated trajectory synthesis framework: (1) crawling open-source online tutorials and parsing them into structured, multi-step task specifications; (2) orchestrating a vision-language model (VLM) agent to execute tasks and generate trajectories in real GUI environments; and (3) employing a VLM-based evaluator for end-to-end automatic trajectory validation. We introduce “guided replay”—the first paradigm enabling fully automated conversion of unstructured textual tutorials into executable, verifiable GUI trajectories without human annotation. Contribution/Results: Experiments demonstrate that synthesized trajectories significantly improve agent performance in GUI element localization and multi-step planning, outperforming prior methods across multiple benchmarks. Moreover, the per-trajectory data cost is reduced by over an order of magnitude, enabling scalable, low-cost GUI agent training.
This work addresses privacy-preserving anomaly detection in human mobility by generating synthetic trajectories that preserve spatiotemporal fidelity while resisting machine learning–based re-identification. Method: To overcome limitations in interpretability and efficiency of existing approaches, we propose a novel trajectory generation framework integrating abductive reasoning with A* search—guided by a parsimony function grounded in aggregated ground truth—and a subset lower-bound estimation mechanism to ensure computational efficiency and attributional explainability. The method unifies annotated logic programming, bottom-up rule learning, and geographic knowledge graph retrieval, and is designed for cloud-native deployment. Contribution/Results: Experimental evaluation demonstrates high fidelity of generated trajectories even at ultra-large scale; the approach has undergone government field validation and is operationally deployed in real-world security systems.
Existing datasets of scientific ideation trajectories struggle to comprehensively capture the full research process—from literature exploration and tool utilization to the evolution of intermediate artifacts and final proposals. This work proposes a reverse-to-forward synthesis mechanism that emulates the uncertainty, evidence integration, and phased convergence characteristic of real scientific inquiry through a Generator–Advisor architecture. By leveraging action–observation–editing sequence modeling, context-aware verification, and process-level supervision, the approach generates multi-turn trajectories aligned with authentic research practices, starting from high-quality papers and proposals. The study yields the first trajectory dataset spanning the complete scientific workflow and establishes a generalizable paradigm for synthesizing process-supervised data for scientific agents.
This study addresses the limitations of real-world human mobility data, which are often sparse and subject to participant bias, as well as the shortcomings of existing synthetic generation methods that struggle to balance realism with controllability. To overcome these challenges, this work proposes an integrated framework that unifies OpenStreetMap-based geographic simulation, genetic algorithm–driven parameter calibration, Patterns-of-Life behavioral modeling, structured data processing, and visualization analytics. The framework enables the generation of high-fidelity, scalable, individual-level synthetic trajectories. Empirical evaluation demonstrates that the resulting large-scale synthetic dataset closely matches real-world data in key statistical characteristics—such as daily trip frequency and activity radius—thereby providing a robust foundation for downstream modeling tasks and benchmarking studies.
To address rare, transient, and precisely localized dynamic scientific phenomena—such as volcanic eruption plumes—this paper proposes an onboard real-time perception–decision–response closed-loop framework. Methodologically, it fuses forward-looking satellite imagery with lightweight CNNs and traditional machine learning models for edge-based plume detection, and integrates multi-objective trajectory planning to autonomously generate optimal high-resolution sensor pointing paths. Its key contribution lies in the first deep integration of event-driven real-time detection and online trajectory planning on an edge computing platform, enabling fully autonomous onboard observation scheduling. Simulation results demonstrate that, compared to baseline approaches, the framework achieves over a tenfold increase in scientific return, while total inference and planning latency remains below typical revisit intervals—substantially improving capture probability and data value for sparse dynamic events.