REALM: A Coarse-to-Fine Generative Framework for Embodied Reactive Listening

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of temporal alignment and modeling local expression stochasticity in listener facial motion generation for embodied conversational AI. To this end, it proposes the REALM framework, which employs a coarse-to-fine generation strategy integrating a delay-centered attention prior with an adaptive gating mechanism to balance global trajectories and local variations. Furthermore, a reactive gated speaker-listener fusion module and an audio-conditioned stochastic residual enhancement technique are introduced to achieve high-quality, audio-driven responsive listening. Extensive evaluations on the ViCo and L2L datasets demonstrate that REALM outperforms existing baselines across multiple metrics. The practical effectiveness of the proposed approach is further validated through real-world deployment on the Ameca robot and corroborating user studies.
📝 Abstract
Generating responsive listener facial motion is an important task for embodied conversational AI. Two modeling challenges are central: accounting for the timing of speaker cues while maintaining continuity with the listener's ongoing motion, and capturing locally variable facial events alongside the overall motion trajectory. Listener responses may follow preceding cues with a temporal lag, while brief expressions and blinks introduce variation that is difficult to predict deterministically. These challenges motivate a framework that combines history-aware temporal alignment with stochastic expression refinement. We propose REALM (Reactive Embodied Audio-driven Listening Model), a coarse-to-fine framework for audio-driven reactive listening. A Reactive Gated Speaker-Listener Fusion module combines listener motion history with speaker audio through a delay-centered attention prior and adaptive gating. A coarse decoder predicts a base motion trajectory, which is augmented by audio-conditioned stochastic residuals in the expression subspace while retaining the coarse pose parameters. Evaluations on ViCo and L2L show improvements over the evaluated baselines across multiple motion-quality metrics. Additional analyses examine delay sensitivity, gate behavior, and blink dynamics. Finally, deployment on an Ameca humanoid robot and a perceptual user study demonstrate the applicability of the generated behavior to physical embodiment. Code: https://github.com/lipzh5/REALM Demo: https://youtu.be/Tf5mpd5S8VQ
Problem

Research questions and friction points this paper is trying to address.

Embodied Conversational AI
Reactive Listening
Listener Facial Motion Generation
Temporal Alignment
Stochastic Expression
Innovation

Methods, ideas, or system contributions that make the work stand out.

Coarse-to-Fine Generation
Reactive Listening
Gated Speaker-Listener Fusion
Stochastic Expression Refinement
Embodied AI
🔎 Similar Papers
No similar papers found.