iMotion-LLM: Motion Prediction Instruction Tuning

๐Ÿ“… 2024-06-10
๐Ÿ›๏ธ arXiv.org
๐Ÿ“ˆ Citations: 4
โœจ Influential: 1
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the critical challenge of generating safe, feasible, and context-aware interactive motion trajectories for autonomous driving conditioned on natural language instructions. Methodologically, we propose a semantic-guided multimodal motion prediction framework centered on the first text-instruction-driven multimodal large language model (MLLM), integrating a pretrained LLM, LoRA-based efficient fine-tuning, and multimodal scene encodingโ€”trained on our newly curated InstructWaymo dataset. Crucially, we introduce the first instruction feasibility identification module with an active rejection mechanism to handle infeasible commands. Experiments on the Waymo Open Motion Dataset demonstrate that our model achieves high trajectory generation accuracy for feasible instructions and significantly outperforms baselines in rejecting infeasible ones. These results validate the effectiveness of semantic guidance in enhancing dynamic scene understanding and safety-critical response capabilities.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Safety and RobustnessPlanning, Routing, and Scheduling: Planning with Language Models

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
๐Ÿ“ Abstract
We introduce iMotion-LLM: a Multimodal Large Language Models (LLMs) with trajectory prediction, tailored to guide interactive multi-agent scenarios. Different from conventional motion prediction approaches, iMotion-LLM capitalizes on textual instructions as key inputs for generating contextually relevant trajectories. By enriching the real-world driving scenarios in the Waymo Open Dataset with textual motion instructions, we created InstructWaymo. Leveraging this dataset, iMotion-LLM integrates a pretrained LLM, fine-tuned with LoRA, to translate scene features into the LLM input space. iMotion-LLM offers significant advantages over conventional motion prediction models. First, it can generate trajectories that align with the provided instructions if it is a feasible direction. Second, when given an infeasible direction, it can reject the instruction, thereby enhancing safety. These findings act as milestones in empowering autonomous navigation systems to interpret and predict the dynamics of multi-agent environments, laying the groundwork for future advancements in this field.
Problem

Research questions and friction points this paper is trying to address.

Generates safe, instruction-based trajectories for autonomous driving
Combines LLM with trajectory prediction for interpretable motion generation
Enables text-guided adaptable driving behavior using multimodal datasets
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM integrated with trajectory prediction modules
Generates trajectories from textual instructions
Uses LoRA fine-tuning and multimodal encoder-decoder
๐Ÿ”Ž Similar Papers
๐Ÿ’ผ Related Jobs
No related jobs found.