avatar animation

Designing and implementing embodied conversational agents with synchronized speech and animation to present findings, support user sensemaking, and foster therapeutic rapport through coordinated multimodal behaviors.

avataranimation

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body

Dec 16, 2025
JZ
Juze Zhang
🏛️ Stanford University | ByteDance

Existing co-speech gesture generation systems predominantly model language-to-motion translation as a static mapping task, lacking joint decision-making capabilities regarding *when* to move, *how* to move, and *how to dynamically adapt* gestures to conversational context—leading to temporal fragility, weak social interactivity, and disintegration across speech, text, and motion modalities. This paper introduces the first Speech-Language-Behavior (SLB) unified modeling framework. We propose a Modality-Partitioned Mixture-of-Experts (MoME) architecture that enables multimodal synchronous planning via cross-modal hard routing and cross-attention. Furthermore, we integrate streaming behavior control hooks and multimodal token interleaving to support end-to-end co-generation and hybrid-initiated interaction of speech, facial expressions, and body gestures. Evaluations on multi-turn dialogues demonstrate significant improvements in gesture–language alignment and behavioral naturalness, comprehensively outperforming state-of-the-art co-speech gesture and text-driven baselines.

Addresses fragmented modeling of speech, text, and motion in prior systemsDevelops a conversational 3D agent that jointly plans language and movementEnables controllable, socially competent multimodal interaction in dialogue

Existing systems struggle to unify multimodal perception, embodied expression, and multi-agent collaborative decision-making within a shared physical space, limiting natural and scalable human–multi-robot interaction. This work proposes a unified framework for human–multi-agent interaction that integrates multimodal perception, large language model (LLM)-driven embodied planning, and a centralized coordination mechanism within a multi-agent architecture. The mechanism dynamically manages speaking turns and behavioral participation to effectively prevent conflicts and enable coordinated strategies across speech, gesture, gaze, and locomotion. Evaluated on a dual-humanoid robot platform, the system demonstrates robust cross-agent collaborative reasoning and embodied responsiveness, significantly enhancing the naturalness and scalability of human–robot interaction.

coordinated decision-makingembodied expressionhuman-multi-agent interaction

Existing open dialogue datasets for mental health lack authenticity and diversity, while the implicit nature of psychological processes—particularly in clients—hampers fidelity in synthetic dialogue generation. Method: This paper proposes an Embodied Conversational Agent (ECA) simulation framework tailored for psychological counseling, integrating embodied cognition theory with clinical practice and leveraging large language models for memory-augmented generation. It establishes the first theory-guided embodied memory space and a scalable ECA paradigm. Six principled design objectives are introduced, with dialogue modeling driven by high-frequency counseling questions and empirically validated on the D4 dataset. Contribution/Results: Experiments demonstrate that generated dialogues significantly outperform baselines in authenticity and necessity; licensed counselors confirm their professional validity; and a high-quality public ECA dataset has been released.

Addressing authenticity challenges in synthetic mental health dialogue dataImproving diversity of psychological counseling simulations through embodied agentsIncorporating psychological theories into LLM-based conversational agent frameworks

This work addresses the lack of empathy and naturalness in current conversational agents, which stems from their failure to model temporal cues inherent in human active listening—such as contextualized silence. The authors systematically identify and formalize five context-aware pacing strategies employed by humans during active listening, including reflective silence and empathetic silence, and propose a novel dialogue agent capable of dynamically modulating its response timing based on these strategies. Through qualitative analysis, user studies, and controlled experiments across two interaction scenarios, the approach demonstrates significant improvements over fixed-pacing baselines, yielding enhanced perceptions of anthropomorphism, conversational fluency, engagement, depth of self-disclosure, and emotional trust.

active listeningcontext-aware pacingconversational agents

GOMA: Proactive Embodied Cooperative Communication via Goal-Oriented Mental Alignment

Mar 17, 2024
LY
Lance Ying
🏛️ Harvard University | Dartmouth College | Johns Hopkins University | Massachusetts Institute of Technology

In human-robot collaboration, inefficient goal and action alignment arises from information asymmetry between agents. Method: This paper proposes a goal-oriented mental alignment framework enabling embodied AI assistants to proactively initiate natural-language communication aligned with shared task objectives. We formulate spoken interaction as a planning problem that minimizes goal-relevant mental state misalignment, thereby enabling intention-driven, anticipatory language generation. The framework integrates embodied reasoning, goal-conditioned mental modeling, and context-aware language generation—bypassing reliance on end-to-end large language model outputs. Results: Evaluated on the Overcooked and VirtualHome benchmarks, our approach achieves significant improvements in collaborative task success rates. Human user studies further demonstrate substantial gains in perceived assistant trustworthiness and usability, validating the effectiveness of goal-aligned, cognitively grounded communication.

AI AssistanceHuman-AI CoordinationNatural Language Communication

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing conversational agents that employ static personas, which often fail to adapt to dynamic shifts in task context, user goals, and situational urgency, leading to interactional mismatches. To overcome this, the authors propose a fluid persona framework that jointly models metaphorical roles—such as coach or mentor—and the intensity of personality expression (low, medium, high) to enable real-time adaptation based on task context, user traits, and situational urgency. Built upon large language models, the framework integrates context-aware mechanisms with personality dimension modulation strategies, allowing dynamic switching of both role and expression intensity during dialogue. Empirical evaluations demonstrate that this approach significantly enhances user experience, trust, and willingness to adopt recommended behaviors across diverse domains, including medical consultation, fitness coaching, and reflective learning.

behavior changeconversational agentsmetaphorical persona

Traditional mental health services are often scarce and costly, while existing large language model (LLM)-driven conversational systems commonly suffer from insufficient empathy, lack of personalization, and low factual reliability. To address these limitations, this work proposes a lightweight virtual conversational agent framework that integrates retrieval-augmented generation (RAG), structured user memory, and multimodal interaction to enable high-quality, cross-culturally personalized empathetic dialogue—even on smaller-scale models. Experimental results demonstrate significant improvements in retrieval accuracy and response quality on objective metrics. Furthermore, user studies confirm that the system substantially outperforms pure LLM baselines in coherence, factual accuracy, and perceived empathy, with a clear majority of participants expressing a strong preference for the proposed approach.

conversational agentempathyfactual grounding

This study addresses the limitations of traditional wearable health data visualizations, which often rely on static charts and lack intuitive, interactive mechanisms for user reflection. To overcome this, the authors propose a novel “embodied dialogue” paradigm featuring a dual-agent architecture—comprising an Observer and a Presenter—that integrates lightweight data preprocessing with conversational statistical summarization. Implemented in Unity, the embodied dialogue agent translates objective health trends into natural language narratives while deliberately avoiding clinical recommendations. A user study (N=5) demonstrates that, compared to conventional dashboards, this approach significantly enhances users’ comprehension of their data and the concreteness of their intended actions, fostering a cognitive shift from passive data viewing toward active meaning-making.

data reflectionembodied conversationhuman-computer interaction

This study addresses the global shortage of evidence-based psychotherapy resources—even in high-income regions, where long wait times persist—by proposing a large language model–based embodied conversational agent. The system uniquely integrates multilevel psychological analysis, process-oriented therapeutic principles, and embodied interaction, leveraging retrieval-augmented generation, emotion recognition, psychological flexibility assessment, and synchronized speech animation to deliver real-time, clinically safe, and evidence-informed responses. Evaluated under a GPT-5.2 configuration, the agent outperformed human therapist responses in comprehension, interpersonal effectiveness, collaboration, and therapeutic adherence, and received endorsement from eleven licensed psychotherapists. This work establishes a novel, supervisable, and safety-controlled paradigm for AI-delivered psychological support.

evidence-based therapymental health supportpsychotherapy access

This work proposes an emotion-aware virtual reality (VR) interaction system that addresses the limitation of existing VR conversational agents, which predominantly rely on textual semantics while neglecting the rich emotional cues embedded in vocal prosody, often resulting in emotionally inconsistent responses. The proposed system explicitly integrates real-time emotion labels—derived from speech prosody analysis and affect recognition—into the dialogue context of a large language model (LLM), thereby enabling emotionally aligned response generation. By unifying vocal prosody analysis, real-time emotion recognition, and LLM-driven dialogue mechanisms, the system significantly enhances perceived dialogue quality, naturalness, engagement, warmth, and human-likeness in user studies involving 30 participants, with 93.3% of users expressing a clear preference for the emotion-aware agent.

conversational agentsemotion recognitionemotional context

Hot Scholars

SS

Shunsuke Saito

Research Scientist, Meta Codec Avatars Lab
Digital HumansComputer VisionComputer Graphics
MH

Marc Habermann

Senior Researcher, Max Planck Institute for Informatics
Computer VisionComputer GraphicsMachine LearningHuman Performance Capture
LB

Liefeng Bo

Head of Applied Computer Vision Lab at Alibaba Group
Machine LearningComputer VisionRobotics
YS

Yujun Shen

Ant Group
Generative ModelingComputer VisionDeep Learning
CT

Christian Theobalt

Professor, Max Planck Institute for Informatics, Saarland Informatics Campus, Saarland University
Computer GraphicsComputer VisionAI & Machine LearningHCI