Automated Evaluation of Multi-Turn Dialogues in In-Car Conversational Assistants

📅 2026-09-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of evaluating in-vehicle assistants during multi-turn interactions, where explicit ground truth is absent and assessing context retention alongside safety constraints remains difficult. To this end, we propose a black-box automated testing framework that constructs a closed-loop simulation environment. The approach incorporates a policy-guided user simulator and an adversarial manager to generate diverse test scenarios, coupled with a dual-layer large language model judge mechanism for automated evaluation. Experimental results demonstrate that the framework’s automated assessments align closely with human judgments. Furthermore, it identifies 2.96 times more failure types and doubles the detection rate of failed sessions. Overall, this work provides an efficient solution for verifying the safety and robustness of in-vehicle dialogue systems.
📝 Abstract
In-car conversational assistants (ICAs) are increasingly integrated into vehicles to support route planning, vehicle control, and information access. Ensuring their reliability is challenging due to multi-turn interactions, the absence of explicit ground truth, and strict safety constraints. Existing evaluation techniques fall short, as they target single-turn settings and fail to capture constraint handling, context retention, and safety-critical behavior across turns. We propose an automated framework for testing the multi-turn conversational capabilities of ICAs. The system is treated as a black box and evaluated via closed-loop simulation with a strategy-guided user simulator, an adversarial strategy manager, and a two-tier LLM judge assessing turn-level failures and conversation-level quality. We evaluate the approach on an industrial ICA with six LLM backends and twelve human annotators. The automated judge shows substantial agreement with humans, and strategy guidance uncovers 2.96 times more unique failure types per conversation and more than doubles the number of unique failing conversations compared to unguided simulation.
Problem

Research questions and friction points this paper is trying to address.

In-car conversational assistants
Multi-turn dialogues
Automated evaluation
Safety constraints
Context retention
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-car conversational assistants
Multi-turn dialogue evaluation
Closed-loop simulation
Adversarial strategy manager
Two-tier LLM judge
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
V
Vaishnav Negi
Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany; BMW Group, Germany
L
Lev Sorokin
Technical University of Munich, Munich, Germany; BMW Group, Germany
Soroosh Tayebi Arasteh
Soroosh Tayebi Arasteh
RWTH Aachen University
Deep LearningAI in MedicineGenerative AIMedical Image Analysis
Andrea Stocco
Andrea Stocco
Technical University of Munich
Software EngineeringSoftware TestingTest AutomationDeep Learning TestingWeb Testing