doPlan: A Variable-Horizon Dataset for Multi-Stage Language-Conditioned Planning in Autonomous Driving

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of long-horizon intent understanding for language instructions in autonomous driving and the limitation of existing datasets to short-term interactions. Leveraging the nuPlan simulation platform, we construct the first real-world, multi-stage, language-conditioned planning dataset. By designing variable time windows and a manual annotation interface, we generate 5,154 driving instructions encompassing intents ranging from immediate to persistent. Experimental evaluations of four language-driven models demonstrate that most associated maneuvers extend beyond a five-second prediction horizon, revealing significant deficiencies in current methods regarding long-horizon intent responsiveness. These findings validate the necessity and research value of the proposed dataset for advancing language-guided autonomous driving.
📝 Abstract
Autonomous vehicles interacting with passengers through natural language must reason beyond immediate commands. Passenger intent may span multiple stages of behavior, depend on future events, refer to surrounding agents or landmarks, and remain relevant as driving conditions evolve. Existing language-enabled driving datasets largely focus on short, localized interactions, leaving these longer-horizon forms of passenger intent comparatively underexplored. We introduce doPlan, to our knowledge the first publicly available, human-annotated real-world dataset designed to study passenger language as persistent task context. Built on nuPlan, doPlan contains 5,154 human-written passenger instructions spanning 169.1 hours of cumulative instruction-aligned context over 50.9 hours of unique driving, with annotation windows ranging from 30.0 to 508.8 s. The annotations capture immediate, deferred, event-conditioned, persistent, and multi-stage passenger intent. The dataset, annotation interface, and supporting resources are publicly available at https://github.com/Mi3-Lab/doPlan. We evaluate four language-conditioned driving models and find that sensitivity to passenger language does not reliably translate into behavior consistent with the requested direction. More broadly, among 2,161 examples with a matched future maneuver, the first associated maneuver occurs a median of 24.6 s after the evaluation point, and only 9.8% occur within the models' common 5 s prediction horizon. These findings highlight the need to connect persistent passenger intent with successive planning decisions. doPlan provides a setting for studying how unresolved goals can be retained, grounded in evolving scenes, and tracked across multiple stages, including how a planner determines when a future goal becomes relevant to the current plan.
Problem

Research questions and friction points this paper is trying to address.

Autonomous Driving
Language-Conditioned Planning
Multi-Stage Passenger Intent
Long-Horizon Planning
Natural Language Interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Language-Conditioned Planning
Variable-Horizon Dataset
Multi-Stage Intent
Autonomous Driving
Persistent Task Context
P
Parthib Roy
Machine Intelligence, Interaction, and Imagination (Mi3) Laboratory, University of California, Merced, Merced, CA, USA.
Y
Yash Tandon
Machine Intelligence, Interaction, and Imagination (Mi3) Laboratory, University of California, Merced, Merced, CA, USA.; Laboratory for Intelligent & Safe Automobiles (LISA), University of California, San Diego, La Jolla, CA, USA.
M
Marcus Blennemann
Machine Intelligence, Interaction, and Imagination (Mi3) Laboratory, University of California, Merced, Merced, CA, USA.; Laboratory for Intelligent & Safe Automobiles (LISA), University of California, San Diego, La Jolla, CA, USA.
G
Giovanni Tapia Lopez
Machine Intelligence, Interaction, and Imagination (Mi3) Laboratory, University of California, Merced, Merced, CA, USA.
A
Angel Martinez-Sanchez
Machine Intelligence, Interaction, and Imagination (Mi3) Laboratory, University of California, Merced, Merced, CA, USA.
M
Mohan M. Trivedi
Laboratory for Intelligent & Safe Automobiles (LISA), University of California, San Diego, La Jolla, CA, USA.
Ross Greer
Ross Greer
University of California Merced
Artificial IntelligenceMachine VisionAutonomous DrivingHuman-Robot InteractionComputer Music