DriveHierarchy: A Benchmark for Diagnosing VLM Driving Capabilities from Open-Loop Understanding to Closed-Loop Execution

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of structured hierarchies and open-loop–closed-loop diagnostic correlations in evaluating Vision-Language Models (VLMs) for autonomous driving. To this end, it proposes a four-level capability architecture—encompassing perception, memory, reasoning, and execution—to construct a unified hierarchical benchmark. Methodologically, multi-source data are fused to generate large-scale question-answering pairs, and an interactive closed-loop simulation platform grounded in real-world road networks is developed to establish mappings between open-loop comprehension and closed-loop execution. Experiments systematically evaluate fifteen VLMs, revealing significant disparities in their structured, non-redundant capabilities. These findings enable fine-grained model diagnosis and provide a practical foundation for benchmark-guided optimization of VLM-based driving systems.
📝 Abstract
Evaluating VLM-based autonomous driving remains difficult because driving competence is composite, where a capable system must ground traffic participants and hazards, integrate context across views and time, reason about future evolution, and act appropriately under closed-loop interaction. Existing benchmarks usually assess either open-loop understanding or closed-loop driving but provide limited structure for explaining how these abilities are organized, how they relate, and how they may inform model diagnosis and improvement. We present \textsc{DriveHierarchy}, a hierarchical benchmark that organizes VLM-based autonomous driving into four ranks, spanning perceptual grounding, contextual memory, mental reasoning, and closed-loop execution. To instantiate this hierarchy, we integrate multiple open-source autonomous-driving datasets into a unified open-loop benchmark with 76,798 question-answer pairs over 84,279 frames and develop a closed-loop simulation platform with interactive scenario construction on a real-world road network, from which 100 driving scenarios are curated for embodied evaluation. Experiments on 15 VLMs show that \textsc{DriveHierarchy} captures structured but non-redundant capability variation, relates open-loop understanding to closed-loop driving, and provides a practical basis for diagnosis and benchmark-guided optimization. \textsc{DriveHierarchy} therefore serves as a unified framework for evaluating and improving VLM-based autonomous driving systems. An anonymized project has been released on https://github.com/PerfectXu88/DriveHierarchy
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Autonomous Driving
Benchmark Evaluation
Open-loop Understanding
Closed-loop Execution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Benchmark
Vision-Language Models
Autonomous Driving
Closed-Loop Simulation
Capability Diagnosis
🔎 Similar Papers
No similar papers found.