CausalDriveBench: Evaluating Causal Reasoning in Vision-Language-Action Models for Autonomous Driving
This study investigates whether Vision-Language-Action (VLA) models for autonomous driving possess genuine causal reasoning capabilities. Grounded in Pearl’s hierarchy of causation, we construct a four-tier evaluation framework spanning association, intervention, and counterfactuals. By distinguishing perceptual salience from causal relevance through causal scene graphs, structured visual question answering, and counterfactual trajectory generation, this framework enables dual verification of reasoning and action on the nuScenes dataset. Experimental results reveal that even state-of-the-art models achieve only 70.6% accuracy, demonstrating that fluent reasoning does not equate to causal understanding. Furthermore, fine-tuning is shown to potentially degrade these causal capabilities. This work establishes a critical benchmark and offers new perspectives for evaluating and enhancing VLA models in autonomous driving.