🤖 AI Summary
This study investigates the impact of virtual patient facial expression intensity on perceived realism and emotional intelligibility during the delivery of bad news in medical communication training. The authors developed an AI-driven virtual patient system integrating large language model–powered dialogue with real-time VR facial animation, and conducted a formative study combining expert qualitative interviews and quantitative evaluation. Their findings reveal, for the first time, that facial expression intensity must be coordinated with vocal prosody and bodily posture as part of a cohesive multimodal ensemble to achieve holistic emotional realism; modulating facial expressions in isolation yields limited effectiveness. Based on these insights, the study proposes a synchronization design roadmap for multimodal emotional expression in medical simulators, offering both theoretical grounding and practical guidance for developing high-fidelity virtual patients.
📝 Abstract
Interactive virtual patients driven by large language models (LLMs) offer scalable solutions for medical communication training, such as breaking bad news. However, designing their emotional expressiveness remains a challenge. This paper presents an AI-driven virtual patient framework combining LLM dialogue with real-time facial animation in virtual reality (VR). We conducted an exploratory, formative evaluation with seven medical experts to gather early feedback and elicit design requirements. The evaluation focused on how variations in facial expression intensity affect perceived realism and the virtual patient's emotion intelligibility. While descriptive quantitative ratings remained baseline across conditions, qualitative interviews provided deep insights into how experts perceive virtual emotional cues. The findings suggest that experts evaluate emotional realism holistically through multiple verbal and non-verbal channels; isolated facial adjustments are easily overshadowed by dialogue nuances and vocal prosody. Based on these insights, we present a roadmap for future medical training simulators, highlighting the need for synchronized, multi-modal pipelines incorporating body gestures, conversational pauses, and vocal dynamics.