🤖 AI Summary
This work proposes an emotion-aware virtual reality (VR) interaction system that addresses the limitation of existing VR conversational agents, which predominantly rely on textual semantics while neglecting the rich emotional cues embedded in vocal prosody, often resulting in emotionally inconsistent responses. The proposed system explicitly integrates real-time emotion labels—derived from speech prosody analysis and affect recognition—into the dialogue context of a large language model (LLM), thereby enabling emotionally aligned response generation. By unifying vocal prosody analysis, real-time emotion recognition, and LLM-driven dialogue mechanisms, the system significantly enhances perceived dialogue quality, naturalness, engagement, warmth, and human-likeness in user studies involving 30 participants, with 93.3% of users expressing a clear preference for the emotion-aware agent.
📝 Abstract
In VR interactions with embodied conversational agents, users' emotional intent is often conveyed more by how something is said than by what is said. However, most VR agent pipelines rely on speech-to-text processing, discarding prosodic cues and often producing emotionally incongruent responses despite correct semantics. We propose an emotion-context-aware VR interaction pipeline that treats vocal emotion as explicit dialogue context in an LLM-based conversational agent. A real-time speech emotion recognition model infers users' emotional states from prosody, and the resulting emotion labels are injected into the agent's dialogue context to shape response tone and style. Results from a within-subjects VR study (N=30) show significant improvements in dialogue quality, naturalness, engagement, rapport, and human-likeness, with 93.3% of participants preferring the emotion-aware agent.