Score
Designing and implementing embodied conversational agents with synchronized speech and animation to present findings, support user sensemaking, and foster therapeutic rapport through coordinated multimodal behaviors.
Existing co-speech gesture generation systems predominantly model language-to-motion translation as a static mapping task, lacking joint decision-making capabilities regarding *when* to move, *how* to move, and *how to dynamically adapt* gestures to conversational context—leading to temporal fragility, weak social interactivity, and disintegration across speech, text, and motion modalities. This paper introduces the first Speech-Language-Behavior (SLB) unified modeling framework. We propose a Modality-Partitioned Mixture-of-Experts (MoME) architecture that enables multimodal synchronous planning via cross-modal hard routing and cross-attention. Furthermore, we integrate streaming behavior control hooks and multimodal token interleaving to support end-to-end co-generation and hybrid-initiated interaction of speech, facial expressions, and body gestures. Evaluations on multi-turn dialogues demonstrate significant improvements in gesture–language alignment and behavioral naturalness, comprehensively outperforming state-of-the-art co-speech gesture and text-driven baselines.
Existing systems struggle to unify multimodal perception, embodied expression, and multi-agent collaborative decision-making within a shared physical space, limiting natural and scalable human–multi-robot interaction. This work proposes a unified framework for human–multi-agent interaction that integrates multimodal perception, large language model (LLM)-driven embodied planning, and a centralized coordination mechanism within a multi-agent architecture. The mechanism dynamically manages speaking turns and behavioral participation to effectively prevent conflicts and enable coordinated strategies across speech, gesture, gaze, and locomotion. Evaluated on a dual-humanoid robot platform, the system demonstrates robust cross-agent collaborative reasoning and embodied responsiveness, significantly enhancing the naturalness and scalability of human–robot interaction.
Existing open dialogue datasets for mental health lack authenticity and diversity, while the implicit nature of psychological processes—particularly in clients—hampers fidelity in synthetic dialogue generation. Method: This paper proposes an Embodied Conversational Agent (ECA) simulation framework tailored for psychological counseling, integrating embodied cognition theory with clinical practice and leveraging large language models for memory-augmented generation. It establishes the first theory-guided embodied memory space and a scalable ECA paradigm. Six principled design objectives are introduced, with dialogue modeling driven by high-frequency counseling questions and empirically validated on the D4 dataset. Contribution/Results: Experiments demonstrate that generated dialogues significantly outperform baselines in authenticity and necessity; licensed counselors confirm their professional validity; and a high-quality public ECA dataset has been released.
This work addresses the lack of empathy and naturalness in current conversational agents, which stems from their failure to model temporal cues inherent in human active listening—such as contextualized silence. The authors systematically identify and formalize five context-aware pacing strategies employed by humans during active listening, including reflective silence and empathetic silence, and propose a novel dialogue agent capable of dynamically modulating its response timing based on these strategies. Through qualitative analysis, user studies, and controlled experiments across two interaction scenarios, the approach demonstrates significant improvements over fixed-pacing baselines, yielding enhanced perceptions of anthropomorphism, conversational fluency, engagement, depth of self-disclosure, and emotional trust.
In human-robot collaboration, inefficient goal and action alignment arises from information asymmetry between agents. Method: This paper proposes a goal-oriented mental alignment framework enabling embodied AI assistants to proactively initiate natural-language communication aligned with shared task objectives. We formulate spoken interaction as a planning problem that minimizes goal-relevant mental state misalignment, thereby enabling intention-driven, anticipatory language generation. The framework integrates embodied reasoning, goal-conditioned mental modeling, and context-aware language generation—bypassing reliance on end-to-end large language model outputs. Results: Evaluated on the Overcooked and VirtualHome benchmarks, our approach achieves significant improvements in collaborative task success rates. Human user studies further demonstrate substantial gains in perceived assistant trustworthiness and usability, validating the effectiveness of goal-aligned, cognitively grounded communication.
This work addresses the limitations of existing conversational agents that employ static personas, which often fail to adapt to dynamic shifts in task context, user goals, and situational urgency, leading to interactional mismatches. To overcome this, the authors propose a fluid persona framework that jointly models metaphorical roles—such as coach or mentor—and the intensity of personality expression (low, medium, high) to enable real-time adaptation based on task context, user traits, and situational urgency. Built upon large language models, the framework integrates context-aware mechanisms with personality dimension modulation strategies, allowing dynamic switching of both role and expression intensity during dialogue. Empirical evaluations demonstrate that this approach significantly enhances user experience, trust, and willingness to adopt recommended behaviors across diverse domains, including medical consultation, fitness coaching, and reflective learning.
Traditional mental health services are often scarce and costly, while existing large language model (LLM)-driven conversational systems commonly suffer from insufficient empathy, lack of personalization, and low factual reliability. To address these limitations, this work proposes a lightweight virtual conversational agent framework that integrates retrieval-augmented generation (RAG), structured user memory, and multimodal interaction to enable high-quality, cross-culturally personalized empathetic dialogue—even on smaller-scale models. Experimental results demonstrate significant improvements in retrieval accuracy and response quality on objective metrics. Furthermore, user studies confirm that the system substantially outperforms pure LLM baselines in coherence, factual accuracy, and perceived empathy, with a clear majority of participants expressing a strong preference for the proposed approach.
This study addresses the limitations of traditional wearable health data visualizations, which often rely on static charts and lack intuitive, interactive mechanisms for user reflection. To overcome this, the authors propose a novel “embodied dialogue” paradigm featuring a dual-agent architecture—comprising an Observer and a Presenter—that integrates lightweight data preprocessing with conversational statistical summarization. Implemented in Unity, the embodied dialogue agent translates objective health trends into natural language narratives while deliberately avoiding clinical recommendations. A user study (N=5) demonstrates that, compared to conventional dashboards, this approach significantly enhances users’ comprehension of their data and the concreteness of their intended actions, fostering a cognitive shift from passive data viewing toward active meaning-making.
This study addresses the global shortage of evidence-based psychotherapy resources—even in high-income regions, where long wait times persist—by proposing a large language model–based embodied conversational agent. The system uniquely integrates multilevel psychological analysis, process-oriented therapeutic principles, and embodied interaction, leveraging retrieval-augmented generation, emotion recognition, psychological flexibility assessment, and synchronized speech animation to deliver real-time, clinically safe, and evidence-informed responses. Evaluated under a GPT-5.2 configuration, the agent outperformed human therapist responses in comprehension, interpersonal effectiveness, collaboration, and therapeutic adherence, and received endorsement from eleven licensed psychotherapists. This work establishes a novel, supervisable, and safety-controlled paradigm for AI-delivered psychological support.
This work proposes an emotion-aware virtual reality (VR) interaction system that addresses the limitation of existing VR conversational agents, which predominantly rely on textual semantics while neglecting the rich emotional cues embedded in vocal prosody, often resulting in emotionally inconsistent responses. The proposed system explicitly integrates real-time emotion labels—derived from speech prosody analysis and affect recognition—into the dialogue context of a large language model (LLM), thereby enabling emotionally aligned response generation. By unifying vocal prosody analysis, real-time emotion recognition, and LLM-driven dialogue mechanisms, the system significantly enhances perceived dialogue quality, naturalness, engagement, warmth, and human-likeness in user studies involving 30 participants, with 93.3% of users expressing a clear preference for the emotion-aware agent.