Towards Reliable Vision-Language Models for Autonomous Driving

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of vision-language models (VLMs) in autonomous driving to visual degradations, which lead to inaccurate predictions and unreliable confidence estimates. Leveraging a driving question-answering dataset, this work systematically evaluates the robustness of five VLMs under diverse visual corruptions, quantifying their differential impacts on accuracy and confidence calibration while validating the effectiveness of a visual evidence augmentation (VEA) strategy at inference time. The findings reveal specific interference mechanisms through which visual degradation undermines model reliability. Experimental results demonstrate that although VEA improves the performance of certain VLMs, these gains exhibit significant inconsistency across driving scenarios. Ultimately, this research provides critical empirical evidence for developing more reliable multimodal systems in autonomous driving applications.
📝 Abstract
Vision-Language models (VLMs) are increasingly being explored in autonomous driving for tasks such as scene understanding, driving reasoning, decision-making, and end-to-end driving. As their role becomes more prominent, ensuring their robustness and reliability is increasingly important. In real-world conditions, visual inputs may be degraded by sensor imperfections and environmental conditions, potentially affecting both model predictions and their associated confidence. Such degradation is especially concerning in autonomous driving, where safety-critical decisions require models to make accurate predictions and recognize when their predictions may be unreliable. In this work, we evaluate five VLMs (Qwen3.5-9B, Gemma4-E4B, LLaVA-OneVision-7B, DriveFusion/DriveFusionQA-4B, and NVIDIA Alpamayo-1.5-10B) across four driving-related QA datasets with different visual input settings, including single-frame, multi-view, multi-frame, and monocular inputs. Our results show that the effects of visual corruption vary across models, datasets, and input settings, with changes in accuracy and confidence reliability and also differing across conditions. We then apply Visual Evidence Augmentation ($\mathrm{V}{\scriptstyle \mathrm{EA}}$), a recent inference-time method to examine whether it can improve model reliability under degraded visual conditions. We find that $\mathrm{V}{\scriptstyle \mathrm{EA}}$ improves performance for some models and datasets, although the gains are not consistent across all settings.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Autonomous Driving
Robustness
Visual Degradation
Reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Autonomous Driving
Visual Corruption
Visual Evidence Augmentation
Reliability
🔎 Similar Papers
No similar papers found.