AutoTraces: Autoregressive Trajectory Forecasting via Multimodal Large Language Models

📅 2026-03-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of long-term human trajectory prediction in densely populated environments, where complex social behaviors render accurate forecasting difficult. The authors propose an autoregressive vision–language–trajectory model that integrates physical coordinates into a multimodal large language model via a novel trajectory point tokenization scheme. A lightweight encoder–decoder architecture enables end-to-end generation of future trajectories. Crucially, the method introduces an automatic chain-of-thought mechanism that infers spatiotemporal interaction patterns directly from visual and trajectory data without requiring manual annotations. The framework supports flexible output lengths and cross-scenario prediction, achieving state-of-the-art performance on long-horizon trajectory forecasting benchmarks while significantly enhancing generalization across diverse environments.

Technology Category

Humans and AI: Human-Aware Planning and Behavior PredictionMachine Learning: Large Multimodal Models (LMMs)Planning, Routing, and Scheduling: Planning with Language Models

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSocial Networks and Social Media: Generative AI / large language models and their impact on social systems
📝 Abstract
We present AutoTraces, an autoregressive vision-language-trajectory model for robot trajectory forecasting in humam-populated environments, which harnesses the inherent reasoning capabilities of large language models (LLMs) to model complex human behaviors. In contrast to prior works that rely solely on textual representations, our key innovation lies in a novel trajectory tokenization scheme, which represents waypoints with point tokens as categorical and positional markers while encoding waypoint numerical values as corresponding point embeddings, seamlessly integrated into the LLM's space through a lightweight encoder-decoder architecture. This design preserves the LLM's native autoregressive generation mechanism while extending it to physical coordinate spaces, facilitates modeling of long-term interactions in trajectory data. We further introduce an automated chain-of-thought (CoT) generation mechanism that leverages a multimodal LLM to infer spatio-temporal relationships from visual observations and trajectory data, eliminating reliance on manual annotation. Through a two-stage training strategy, our AutoTraces achieves SOTA forecasting accuracy, particularly in long-horizon prediction, while exhibiting strong cross-scene generalization and supporting flexible-length forecasting.
Problem

Research questions and friction points this paper is trying to address.

trajectory forecasting
human-robot interaction
long-horizon prediction
multimodal learning
autoregressive modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

trajectory tokenization
autoregressive trajectory forecasting
multimodal large language models
chain-of-thought reasoning
cross-scene generalization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Teng Wang
Teng Wang
Southeast University
Embodied NavigationVisual LocalizationComputer Vision
Y
Yanting Lu
Southeast University
R
Ruize Wang
Southeast University