Perceptually Grounded and Semantics-Aware Evaluation for Holistic Co-Speech Gesture Generation

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the disconnect between objective metrics and human perception in co-speech gesture generation, along with the difficulty of quantifying semantic appropriateness. To this end, it constructs a comprehensive benchmark integrating standardized comparisons, human-centric validation, and fine-grained semantic evaluation. Methodologically, multimodal large language models are leveraged to enhance data annotation, and a Semantic Gesture Preservation (SGP) metric is proposed alongside a perception-aligned composite evaluation framework. Experimental results demonstrate that SGP correlates significantly with human subjective judgments, while the composite metric effectively improves perceptual alignment across all evaluated dimensions. These findings confirm the necessity of calibrating objective metrics against subjective evaluations for more faithful assessment of generated gestures.
📝 Abstract
Holistic and semantics-aware co-speech gesture generation has advanced rapidly, yet evaluation remains behind: objective metrics do not consistently reflect human perception, and semantic appropriateness remains difficult to quantify. We present a perceptually grounded and semantics-aware benchmark that combines standardized model comparison, human-centered metric validation, and fine-grained semantic evaluation. We first curate a list of 13 objective metrics covering different aspects, including distributional similarity, geometric fidelity, kinematic quality, cross-modal synchrony, and semantic appropriateness. For the semantic-appropriateness category, we propose a new metric, Semantic Gesture Preservation (SGP), which measures how far semantic gestures in the ground truth are preserved in the generated gestures. For this, we augment the BEAT2 dataset's annotations using a multi-modal LLM. We then conduct a perceptual study where 101 participants score generated gestures among five dimensions, including human-likeness, motion diversity, absence of animation errors, speech timing and content match. We systematically analyze objective metric--subjective score correlations. Unlike Semantic Score (SC), which shows no significant association with the evaluated perceptual dimensions, SGP is selectively aligned with speech-aware human judgments. We construct five target-specific composite metrics aligned with the subjective dimensions. These composites improve perceptual alignment across all five dimensions, with the largest gains for absence of animation errors and content match, indicating that complementary objective signals can better approximate human judgments than individual metrics alone. Overall, our results show that objective metrics require validation against subjective evaluations.
Problem

Research questions and friction points this paper is trying to address.

co-speech gesture generation
evaluation metrics
perceptual alignment
semantic appropriateness
human perception
Innovation

Methods, ideas, or system contributions that make the work stand out.

Co-Speech Gesture Generation
Semantic Gesture Preservation
Perceptual Evaluation
Multi-modal LLM
Composite Metrics
🔎 Similar Papers
No similar papers found.