When Semantics Matter: Reliability-Aware Semantic-Rhythm Control for Co-Speech Gesture Generation

๐Ÿ“… 2026-09-29
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the over-reliance on acoustic prosody at the expense of textual semantics in co-speech gesture generation by proposing a reliability-aware semantic rhythm control framework. The method introduces a distribution-divergence-based conditional information gain to quantify semantic contributions, employing a dual-branch mechanism that strengthens semantic guidance within content-relevant segments. Furthermore, noise-conditional modulation and latent perturbation techniques are incorporated to enhance robustness, achieving a controllable balance between semantic expression and rhythmic synchronization. Experiments on benchmark datasets demonstrate that the proposed framework effectively reconciles semantic consistency, rhythmic synchrony, and motion diversity, significantly improving the overall reliability of generated gestures.
๐Ÿ“ Abstract
Co-speech gesture generation aims to synthesize natural gestures that are both temporally synchronized with speech and semantically consistent with the spoken content. Although recent methods can generate rhythmically plausible motions, they often rely heavily on acoustic prosody while underutilizing textual semantics, especially when semantic annotations are incomplete, noisy, or unavailable. Consequently, the generated gestures may follow speech rhythm while failing to express the intended semantics. To address this problem, we propose a reliability-aware semantic-rhythm control framework for co-speech gesture generation. We first learn a discrete motion prior that represents continuous gestures in a compact and structured motion-code space. We then introduce a dual-branch semantic contribution estimation mechanism consisting of a full multimodal branch and an audio-only branch. Their distributional discrepancy is formulated as conditional information gain to quantify how much textual semantics changes the predicted motion. Based on this estimate, a controllable semantic-rhythm objective selectively strengthens semantic guidance in content-relevant segments while limiting unnecessary semantic intervention in rhythm-dominant segments. Furthermore, we treat background noise as an acoustic reliability condition and introduce noise-conditioned feature modulation together with beneficial latent perturbation to improve generation robustness under realistic acoustic environments. Experiments on benchmark datasets demonstrate that the proposed framework achieves a favorable balance among semantic expressiveness, rhythmic synchronization, motion diversity, and robustness, enabling reliable and controllable co-speech gesture generation.
Problem

Research questions and friction points this paper is trying to address.

co-speech gesture generation
semantic consistency
rhythmic synchronization
acoustic robustness
textual semantics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Co-speech gesture generation
Semantic-rhythm control
Conditional information gain
Discrete motion prior
Noise-conditioned feature modulation
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
Z
Zhirui Xing
Hainan International College, Communication University of China, Lingshui, China
Long Ye
Long Ye
Communication University of China
Multimedia Signal ProcessingArtificial Intelligence
K
Kaige Li
School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University, Shenzhen, China
Ziyi Xu
Ziyi Xu
ร‰cole Polytechnique Fรฉdรฉrale de Lausanne (EPFL)
Ming Meng
Ming Meng
Dartmouth College