🤖 AI Summary
This study addresses the challenges of temporal alignment and modeling local expression stochasticity in listener facial motion generation for embodied conversational AI. To this end, it proposes the REALM framework, which employs a coarse-to-fine generation strategy integrating a delay-centered attention prior with an adaptive gating mechanism to balance global trajectories and local variations. Furthermore, a reactive gated speaker-listener fusion module and an audio-conditioned stochastic residual enhancement technique are introduced to achieve high-quality, audio-driven responsive listening. Extensive evaluations on the ViCo and L2L datasets demonstrate that REALM outperforms existing baselines across multiple metrics. The practical effectiveness of the proposed approach is further validated through real-world deployment on the Ameca robot and corroborating user studies.
📝 Abstract
Generating responsive listener facial motion is an important task for embodied conversational AI. Two modeling challenges are central: accounting for the timing of speaker cues while maintaining continuity with the listener's ongoing motion, and capturing locally variable facial events alongside the overall motion trajectory. Listener responses may follow preceding cues with a temporal lag, while brief expressions and blinks introduce variation that is difficult to predict deterministically. These challenges motivate a framework that combines history-aware temporal alignment with stochastic expression refinement. We propose REALM (Reactive Embodied Audio-driven Listening Model), a coarse-to-fine framework for audio-driven reactive listening. A Reactive Gated Speaker-Listener Fusion module combines listener motion history with speaker audio through a delay-centered attention prior and adaptive gating. A coarse decoder predicts a base motion trajectory, which is augmented by audio-conditioned stochastic residuals in the expression subspace while retaining the coarse pose parameters. Evaluations on ViCo and L2L show improvements over the evaluated baselines across multiple motion-quality metrics. Additional analyses examine delay sensitivity, gate behavior, and blink dynamics. Finally, deployment on an Ameca humanoid robot and a perceptual user study demonstrate the applicability of the generated behavior to physical embodiment. Code: https://github.com/lipzh5/REALM Demo: https://youtu.be/Tf5mpd5S8VQ