CausalDriveBench: Evaluating Causal Reasoning in Vision-Language-Action Models for Autonomous Driving

πŸ“… 2026-09-26
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study investigates whether Vision-Language-Action (VLA) models for autonomous driving possess genuine causal reasoning capabilities. Grounded in Pearl’s hierarchy of causation, we construct a four-tier evaluation framework spanning association, intervention, and counterfactuals. By distinguishing perceptual salience from causal relevance through causal scene graphs, structured visual question answering, and counterfactual trajectory generation, this framework enables dual verification of reasoning and action on the nuScenes dataset. Experimental results reveal that even state-of-the-art models achieve only 70.6% accuracy, demonstrating that fluent reasoning does not equate to causal understanding. Furthermore, fine-tuning is shown to potentially degrade these causal capabilities. This work establishes a critical benchmark and offers new perspectives for evaluating and enhancing VLA models in autonomous driving.
πŸ“ Abstract
Vision-Language-Action (VLA) models for autonomous driving produce natural-language reasoning alongside predicted trajectories, but whether this reasoning reflects the causal structure of the scene remains untested. We introduce CausalDriveBench, an evaluation framework grounded in Pearl's Causal Hierarchy (PCH) that tests causal reasoning in driving-specific VLAs through structured visual question answering (QA) and alternative-trajectory prediction. To this end, we construct causal scene graphs over nuScenes that distinguish causally active, dormant, and distractor entities, separating perceptual salience from causal relevance. The benchmark spans all four rungs of PCH (association, intervention, and counterfactual along with causal discovery) for QA generation. For the higher rungs, we additionally provide reference trajectories under specified scene modifications, enabling action-level verification that complements reasoning-level evaluation. In total, the benchmark contains 7,285 verified causal QA pairs and 1,000 counterfactual trajectories derived from nuScenes. We evaluate 10 driving-specific VLAs and 3 general-purpose VLMs, and report three findings. First, the best model reaches only 70.6% QA accuracy, and 4 of 13 models score below random chance. Second, comparing each driving VLA to the general-purpose VLM that shares its language backbone, the cost of driving fine-tuning ranges from 2 to 34 percentage points on causal QA, with post-training design explaining the spread. Third, causal QA and trajectory accuracy are statistically uncorrelated across models: under counterfactual prompts, predicted trajectories either over-react or collapse onto the observed-scene baseline. Taken together, these results show that neither fluent rationales nor accurate observed-scene trajectories constitute evidence of causal understanding.
Problem

Research questions and friction points this paper is trying to address.

Causal Reasoning
Vision-Language-Action Models
Autonomous Driving
Evaluation Benchmark
Counterfactual Trajectories
Innovation

Methods, ideas, or system contributions that make the work stand out.

Causal Reasoning
Vision-Language-Action Models
Pearl's Causal Hierarchy
Causal Scene Graphs
Counterfactual Trajectories
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
N
Narendiran Chembu
Fastcode AI
N
Navvrat Rao
Fastcode AI
S
Shreedhar Shreeshail Kodate
Renesas Electronics
G
Gayatri Srujana Banda
Renesas Electronics
A
Arko Sarkar
Renesas Electronics
A
Abhinav Khanna
Indian Institute of Technology Delhi
R
Rajarshee Das
Indian Institute of Technology Delhi
U
Umesh Kanala
Indian Institute of Technology Delhi
S
Siddarth Khandelwal
Fastcode AI
K
Kumar Aman
Fastcode AI
A
Aish Dubey
Renesas Electronics
Kaustubh Beedkar
Kaustubh Beedkar
Indian Institute of Technology Delhi
Arjun Jain
Arjun Jain
Fastcode AI, IISc
Machine LearningComputer VisionComputer Graphics