Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the misalignment between two-dimensional semantic reasoning and three-dimensional geometric planning in Vision-Language-Action (VLA) models for autonomous driving by proposing the GeoCoTDrive framework. Following a "think in 2D, drive in 3D" paradigm, this method constructs an explicit geometry-oriented chain-of-thought for planning by localizing critical regions, retrieving local 3D prior features, and integrating autoregressive context interleaving with region-level ground-truth supervision. Furthermore, the PlanningGrounding dataset is introduced to enhance geometric grounding capabilities. Experimental results demonstrate that the proposed framework significantly improves planning performance in safety-critical scenarios across multiple end-to-end benchmarks, effectively bridging the gap in fine-grained geometric understanding within VLA models.
📝 Abstract
Vision-language-action~(VLA) models have emerged as a promising paradigm for autonomous driving. However, existing VLA models still suffer from a fundamental mismatch: driving actions require precise 3D geometric cues, while visual-language understanding and reasoning are largely conducted in a 2D semantic space. In this paper, we propose GeoCoTDrive, an explicit geometric chain-of-thought framework that grounds geometry in a planning-oriented manner. GeoCoTDrive follows a think with 2D first, drive with dedicated 3D priors paradigm. It first grounds 2D regions corresponding to decision-critical cues, and then retrieves localized 3D priors by sampling features from a geometric foundation model within the grounded regions. These localized geometric features are interleaved into the autoregressive context to support the trajectory generation. To supervise this process, we introduce planning-relevant grounding, a new region-level grounding task that focuses on local spatial cues directly affecting ego planning decisions, and construct the PlanningGrounding dataset to endow VLAs with planning-oriented grounding capability. Experiments across multiple end-to-end autonomous driving benchmarks show that GeoCoTDrive consistently improves safety-critical planning performance, demonstrating the effectiveness of the explicit geometric chain-of-thought process for VLA-based planning.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
Autonomous driving
3D geometric cues
2D semantic space
Fundamental mismatch
Innovation

Methods, ideas, or system contributions that make the work stand out.

Geometric Chain-of-Thought
Vision-Language-Action
Autonomous Driving
3D Priors
Planning-Relevant Grounding