EmoPose: Vision-Language Model Guided Emotion-Aware Gesture Generation for Humanoid Robots

📅 2026-09-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出EmoPose框架,通过视觉-语言模型指导生成情感感知的手势,解决人形机器人在社交互动中表达情感和意图的问题。
📝 Abstract
Socially competent humanoid robots must communicate affect and intent through gesture as well as speech, yet open-ended interaction must become motion that is both expressive and executable on a specific body. This demands semantic flexibility for contextual social intent while preserving deterministic, embodiment-aware robot control. We present EmoPose, a vision-language model (VLM)-guided framework that bridges this gap through an executable semantic interface. Given language, dialogue history, and optional visual context, the VLM selects an ordered gesture plan containing a communicative class, library variant, intensity, and speech anchor. A scalable robot-owned motion library defines the available expressive vocabulary and the source of 14-DoF joint targets. Pose Studio supports automatic trajectory generation, MuJoCo preview, and automatic synchronization of new library entries with the VLM guide; deterministic robot-side modules validate plans, construct trajectories, schedule gestures, and manage queueing and interruption. This division lets the interaction repertoire grow for new social contexts without changing the control interface or delegating raw joint commands to the foundation model. On the EmoPose-Bench, structured GPT-5.5 planning reaches $98.25\pm0.52\%$ on the Easy tier and $76.50\pm0.54\%$ overall, exceeding same-model direct-label prompting. Further tests validate dialogue-context use and ordered multi-action composition. The system completes the nominal MuJoCo suite and realizes all 29 authored variants on the physical Unitree G1. A four-stop laboratory tour demonstrates expressive narration with interruption, camera-grounded dialogue, and navigation.
Problem

Research questions and friction points this paper is trying to address.

humanoid robots
emotion-aware gesture generation
vision-language model
social interaction
expressive communication
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Model
Gesture Generation
Humanoid Robots
Social Interaction
Executable Semantics
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Daojie Peng
The Hong Kong University of Science and Technology (Guangzhou)
B
Bingtao Wang
Shandong University
F
Fulong Ma
The Hong Kong University of Science and Technology (Guangzhou)
W
Wenjun Yue
RoboScience
L
Liang Zhang
Shandong University
J
Jun Ma
The Hong Kong University of Science and Technology (Guangzhou)