Perceptually Grounded and Semantics-Aware Evaluation for Holistic Co-Speech Gesture Generation
This study addresses the disconnect between objective metrics and human perception in co-speech gesture generation, along with the difficulty of quantifying semantic appropriateness. To this end, it constructs a comprehensive benchmark integrating standardized comparisons, human-centric validation, and fine-grained semantic evaluation. Methodologically, multimodal large language models are leveraged to enhance data annotation, and a Semantic Gesture Preservation (SGP) metric is proposed alongside a perception-aligned composite evaluation framework. Experimental results demonstrate that SGP correlates significantly with human subjective judgments, while the composite metric effectively improves perceptual alignment across all evaluated dimensions. These findings confirm the necessity of calibrating objective metrics against subjective evaluations for more faithful assessment of generated gestures.