Beyond Retrieval Relevance: Scene-Grounded Risk Entailment for Vision-Language Driving

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the mismatch between retrieved risk rules and actual driving scenarios in vision-language driving systems, where retrieval relevance does not guarantee scene applicability. To bridge this gap, we propose a framework integrating a driving risk knowledge graph with Semantic Web Rule Language (SWRL) reasoning. Specifically, the method extracts scene facts through structured perception and applies semantic web rules for logical verification, transforming retrieved information into interpretable risk evidence. This approach effectively guides vision-language models in decision-making and assists diffusion planners in generating safe trajectories. Evaluated on the nuReasoning benchmark, the proposed framework improves the planning score by 1.30 and the non-at-fault collision score by 2.76, demonstrating significant enhancements in autonomous driving safety.
📝 Abstract
Retrieval-augmented generation (RAG) gives vision--language driving systems access to external safety knowledge, yet a retrieved risk rule may be relevant without applying to the current scene. A vision--language model (VLM) receiving such knowledge must ground objects, bind entities across time, and verify relations before deciding how to act, leaving the support for risk conclusions implicit. We address this relevance--applicability gap with a Driving-Risk Knowledge Graph (DRKG) and Semantic Web Rule Language (SWRL) reasoning stage before VLM decision-making. Structured perception instantiates scene facts, from which SWRL rules derive events and directed risk relations when their antecedents are jointly satisfied. Recognized events, bound risk relations, and semantic descriptions of activated rules form compact evidence that conditions the VLM and diffusion planner. In matched comparisons on nuReasoning, our method improved the nuReasoning planning score (NPS) by 1.30 points and the non-at-fault collision score (NC) by 2.76 points over the relevance retrieval-based baseline. These gains indicate that scene-applicable risk evidence improves safety-weighted planning relative to semantically retrieved risk knowledge.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Driving
Retrieval-Augmented Generation
Risk Entailment
Scene Grounding
Relevance-Applicability Gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

Driving-Risk Knowledge Graph
SWRL reasoning
Retrieval-Augmented Generation
Vision-Language Model
Diffusion Planner
🔎 Similar Papers
No similar papers found.