Vision-Language Assistant for Emotional Reactions to Risky Driving

📅 2026-07-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the common oversight of driver emotional experience in existing advanced driver assistance systems, which often fail to deliver affectively appropriate feedback while detecting hazardous behaviors. To bridge this gap, the authors propose a vision–language fusion pipeline that integrates YOLOv8-series models for real-time detection of high-risk driving maneuvers—such as sudden lane intrusions—and couples this with emotion-adaptive large language models (e.g., GPT-4o, Claude 3) to generate spoken feedback in varied styles, including neutral, humorous, or analytical tones. A user study (N=108) demonstrates that this approach significantly enhances both safety awareness and perceived comfort among drivers, with the YOLOv8s–GPT-4o configuration achieving the highest rating of 4.29 out of 5.00, thereby validating the efficacy of this emotionally intelligent in-vehicle AI paradigm.
📝 Abstract
This study introduces a vision-language pipeline that detects risky driving behaviors and generates emotionally expressive responses to support driver awareness and comfort. Although vision-language models have advanced perception and reasoning in autonomous driving, existing systems rarely consider the emotional dimension or real-world user experience. Keep Yelling Assistant (KYA) detects high-risk driving maneuvers in real time, such as sudden cut-ins. It then produces emotional responses through a large language model tailored to driver preferences. The framework comprises two core modules. The vision module uses YOLOv8 variants to detect nearby vehicles and identify risky behaviors such as sudden cut-ins. Key driving metrics, including relative distance, speed, and projected reach time, are extracted and normalized to produce a structured behavior log. The language module processes this log with user-defined emotional tone settings, such as neutral, humorous, and analytical, and generates verbal reactions using state-of-the-art large language models, including ChatGPT-4o, Claude 3, Gemini 2.5, and Copilot. We evaluated the proposed system using dashcam videos containing risky driving behaviors and a user study involving 108 participants. Participants selected preferred response styles, and the large language models were evaluated based on emotional alignment. All models received favorable ratings, although preferences varied across personas. Notably, the combination of YOLOv8s and ChatGPT-4o achieved the highest score of 4.29 out of 5.00. By integrating real-world perception with emotionally adaptive dialogue, KYA introduces a new paradigm for emotionally intelligent in-vehicle artificial intelligence. It offers promising directions for improving safety, trust, and emotional well-being in both conventional and autonomous vehicles.
Problem

Research questions and friction points this paper is trying to address.

vision-language models
emotional reactions
risky driving
driver awareness
in-vehicle AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision-language model
emotionally expressive response
risky driving detection
YOLOv8
large language model
🔎 Similar Papers
No similar papers found.