AnchorReasoning: A Visual Grounding and Causal Reasoning Dataset in Long-Tail Autonomous Driving Scenarios

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为了解决长尾自动驾驶场景中视觉证据与决策规划关联不足的问题,本文通过构建包含决策关键元素的AnchorReasoning数据集,并采用分层能力逐步学习策略,提升了视觉定位推理和轨迹预测性能。
📝 Abstract
Vision-language models (VLMs) offer a promising approach to long-tail autonomous driving, but existing driving datasets provide limited supervision for connecting decision-critical visual evidence with reasoning and planning. We introduce AnchorReasoning, a visually grounded reasoning dataset built on WOD-E2E, containing 416,119 annotated frames and 395,379 decision-critical elements across four major categories and 19 fine-grained types. Each frame is organized as a visually grounded chain-of-thought (VG-CoT) that links decision-critical element identification and localization, element attributes and implications, driving-action rationale, and action and trajectory planning. We further develop a curriculum supervised fine-tuning strategy that progressively learns these hierarchical capabilities, together with an object-size-aware grounding metric for evaluating localization quality. Experiments across eight general-purpose, embodied-AI, and AV-specific backbones show that VG-CoT supervision improves grounded reasoning and trajectory prediction. Across models, 5-s ADE and FDE decrease by 7.84 and 11.86, while RFS Frame and Cluster improve by 1.66 and 1.70. These gains are achieved with 18.5 fewer reasoning tokens and 0.32 s/frame lower inference latency on average, demonstrating the value of visually grounded, decision-focused supervision for VLM reasoning and planning in long-tail autonomous driving.
Problem

Research questions and friction points this paper is trying to address.

vision-language models
long-tail autonomous driving
decision-critical visual evidence
reasoning and planning
supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

visually grounded reasoning
curriculum supervised fine-tuning
object-size-aware grounding metric
🔎 Similar Papers
No similar papers found.
Z
Zhipeng Bao
School of Environmental, Civil, Agricultural and Mechanical Engineering, University of Georgia
Wenjie Zhao
Wenjie Zhao
University of Texas at Dallas
computer vision
T
Tianle Zhu
School of Environmental, Civil, Agricultural and Mechanical Engineering, University of Georgia
H
Haohua Que
School of Environmental, Civil, Agricultural and Mechanical Engineering, University of Georgia
C
Chence Yang
School of Computing, University of Georgia
Geng Yuan
Geng Yuan
University of Georgia
Efficient AIExplainable AITrustworthy MLEdge ComputingAI Applications
Q
Qianwen Li
School of Environmental, Civil, Agricultural and Mechanical Engineering, University of Georgia