Less Language, More Latents: Annotation-Efficient VLAs for Driving

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决驾驶模型训练中语言标注稀缺问题,提出LADA方法,通过三阶段流程将无标签观测-轨迹对转化为语言条件控制的基底,显著减少所需语言标注。
📝 Abstract
Vision-language-action models (VLA) promise human-steerable autonomous driving, but their training is bottlenecked by the scarcity of frames paired with natural-language instructions: while camera streams and expert trajectories are logged at scale, language annotations (e.g., turn left at the intersection) remain scarce and expensive to acquire. To address this challenge, we introduce Latent Action Driving Annotations (LADA), a three-stage pipeline that transforms abundant unlabelled observation-trajectory pairs into a substrate for language-conditioned control. First, we train a latent action model with a vector-quantised bottleneck, producing a compact codebook of high-level vehicle intents. Second, a small language-annotated subset is used to train a vision-language translator to map observations and language instructions into this codebook. Third, we train a driving VLA on observation-latent-action pairs over the full unlabelled corpus. Using fewer than 5% of language annotations and without leveraging any auxiliary chain-of-thought reasoning or visual question answering streams, LADA achieves a Driving Score of 87.98 and a Success Rate of 70.46% on the closed-loop Bench2Drive benchmark, matching or surpassing fully supervised baselines.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action Models
Language Annotations Scarcity
Autonomous Driving
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent Action Driving Annotations
vector-quantised bottleneck
vision-language translator
🔎 Similar Papers
No similar papers found.
A
Alexey Zakharov
Robert Bosch GmbH, Germany; Five AI Ltd., United Kingdom
K
Kemal Oksuz
Robert Bosch GmbH, Germany; Five AI Ltd., United Kingdom
Puneet K. Dokania
Puneet K. Dokania
University of Oxford | Bosch (Five AI)
Deep Learning