🤖 AI Summary
This study addresses the disconnect between textual reasoning and visual evidence, as well as the lack of explicit attention focus, in autonomous driving. To overcome these limitations, this work proposes a structured multimodal reasoning framework that introduces explicit visual foci to bridge scene understanding and driving decisions. By precisely aligning object descriptions with corresponding image patches, the proposed approach transcends the constraints of text-only chain-of-thought reasoning. Built upon multimodal large language models, the framework integrates visual focus annotation, gaze prediction, and NAVSIM data processing techniques to achieve end-to-end planning. Experimental results demonstrate that the proposed method outperforms existing baselines in both gaze prediction and planning performance on the W3DA and NAVSIM benchmarks, highlighting its effectiveness in grounding linguistic reasoning within visual contexts for autonomous navigation.
📝 Abstract
Driving decisions depend on both where to focus and how to act on what is seen. Effective driving reasoning must establish which objects matter, where they are, and how they inform the intended action. Text-based rationales can describe a driving response while leaving its correspondence to specific visual evidence implicit. Visual focus provides a concrete starting point for this connection by identifying what matters in the scene and where it is. We propose FocusDrive, a structured multimodal reasoning framework that organizes end-to-end planning around explicit visual focus. It pairs descriptions of decision-relevant objects with image-patch references, bringing explicit visual focus into the reasoning that generates driving plans and trajectories. We first assess this focus representation through driver gaze prediction, then investigate its role in planning reasoning using driving-focus annotations within existing NAVSIM training scenes. Experiments on W3DA and NAVSIM demonstrate competitive gaze prediction and end-to-end planning performance, with FocusDrive improving over text-based chain-of-thought. These results support visual focus as an effective link between scene understanding and driving action.