Semantic-Drive: Democratizing Long-Tail Data Curation via Open-Vocabulary Grounding and Neuro-Symbolic VLM Consensus

πŸ“… 2025-12-12
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Manual screening of long-tail safety events (e.g., jaywalking, construction detours) in autonomous driving video logs is labor-intensive; metadata-based retrieval suffers from low precision; and cloud-based vision-language models (VLMs) face privacy risks and computational bottlenecks. Method: We propose the first localized, privacy-preserving semantic data mining framework. It introduces a novel neuro-symbolic architecture integrating real-time open-vocabulary detection (YOLOE), a lightweight inference-optimized visual-language model, and a β€œSystem 2” multi-model judge-scout consensus mechanism to suppress hallucination and enhance semantic robustness. The framework supports cross-domain alignment across nuScenes and Waymo and runs entirely offline on an RTX 3090. Results: On nuScenes, it achieves an event recall of 0.966β€”104% higher than CLIPβ€”and reduces risk assessment error by 40%, demonstrating superior accuracy, privacy compliance, and practical deployability.

Technology Category

Natural Language Processing: Safety and RobustnessComputer Vision: Visual Reasoning & Symbolic RepresentationsMachine Learning: Neuro-Symbolic Learning

Application Category

Security and Privacy: Data transparency and provenanceSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
πŸ“ Abstract
The development of robust Autonomous Vehicles (AVs) is bottlenecked by the scarcity of "Long-Tail" training data. While fleets collect petabytes of video logs, identifying rare safety-critical events (e.g., erratic jaywalking, construction diversions) remains a manual, cost-prohibitive process. Existing solutions rely on coarse metadata search, which lacks precision, or cloud-based VLMs, which are privacy-invasive and expensive. We introduce Semantic-Drive, a local-first, neuro-symbolic framework for semantic data mining. Our approach decouples perception into two stages: (1) Symbolic Grounding via a real-time open-vocabulary detector (YOLOE) to anchor attention, and (2) Cognitive Analysis via a Reasoning VLM that performs forensic scene analysis. To mitigate hallucination, we implement a "System 2" inference-time alignment strategy, utilizing a multi-model "Judge-Scout" consensus mechanism. Benchmarked on the nuScenes dataset against the Waymo Open Dataset (WOD-E2E) taxonomy, Semantic-Drive achieves a Recall of 0.966 (vs. 0.475 for CLIP) and reduces Risk Assessment Error by 40% compared to single models. The system runs entirely on consumer hardware (NVIDIA RTX 3090), offering a privacy-preserving alternative to the cloud.
Problem

Research questions and friction points this paper is trying to address.

Addresses scarcity of rare safety-critical events in autonomous vehicle training data.
Overcomes limitations of coarse metadata search and privacy-invasive cloud VLMs.
Provides a local, privacy-preserving framework for semantic mining of long-tail data.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Local-first neuro-symbolic framework for semantic data mining
Two-stage perception with open-vocabulary grounding and reasoning VLM
Multi-model consensus mechanism to reduce hallucination and error
πŸ”Ž Similar Papers
No similar papers found.