From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究针对车辆语音命令授权中的安全问题,通过202个场景基准测试和七类决策分类,评估了多种大型语言模型的安全性和决策一致性。
📝 Abstract
Large language models (LLMs) are increasingly integrated into vehicle voice assistants. But linking natural-language requests to vehicle functions creates a safety-critical authorization problem. Before executing a command, the system must choose whether to execute, refuse, clarify, require confirmation, defer to manual control, trigger an emergency response, or make no tool call. To our knowledge, prior evaluations do not isolate this pre-action decision across speaker role, authentication status, vehicle state, and tool availability. We introduce a 202-scenario benchmark with Reference Decisions under a seven-class taxonomy. We evaluate two local open-weight models and three API-based LLMs using Decision Alignment and safety-specific error metrics. Alignment ranges from 40.1% for Llama 3.2 3B to 89.1% for Gemini 3.1 Pro Preview. The API-based models score between 83.2% and 89.1%, with no statistically significant differences among them. Even these models produce two to three False Executes among 161 non-execution scenarios, and persistent errors remain in confirmation and manual-control decisions. A controlled Llama 3.2 3B ablation increases alignment to 40.1% under the structured authorization policy, versus 28.2-29.2% under schema-only and generic-safety baselines, but it does not eliminate False Executes. Structured LLM decisions are therefore insufficient as a standalone safety mechanism, and deployment requires an independent enforcement layer that verifies tool permissions and vehicle-state constraints before invoking any vehicle function.
Problem

Research questions and friction points this paper is trying to address.

LLM Safety
Vehicle Voice Command Authorization
Pre-action Decision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Benchmarking
Vehicle Voice Command Authorization
Safety-Critical Authorization
Decision Alignment
Independent Enforcement Layer
🔎 Similar Papers
No similar papers found.
D
Diba Afroze
University of Louisiana at Lafayette, Lafayette, LA 70503, USA
X
Xingli Zhang
University of Louisiana at Lafayette, Lafayette, LA 70503, USA
Yazhou Tu
Yazhou Tu
Auburn University
cyber-physical system securityside channel analysissensing securityembedded system
X
Xiali Hei
University of Louisiana at Lafayette, Lafayette, LA 70503, USA