🤖 AI Summary
This study addresses the limitations of existing listener reaction generation in dyadic conversations, specifically the lack of fine-grained annotations and evaluation protocols confined to visual realism. To this end, we propose GLARE, an audio-driven framework for listener reaction generation. We first construct the inaugural fine-grained listener dataset encompassing six distinct reaction categories. Subsequently, the framework employs a flow-matching Transformer architecture that integrates Qwen2-Audio prosodic conditioning with a frame-level temporal reaction loss to optimize generation quality. Furthermore, a reaction-oriented multidimensional evaluation protocol is designed. Experimental results demonstrate that the proposed method significantly outperforms existing approaches across multiple metrics, including visual fidelity, reaction accuracy, and temporal alignment.
📝 Abstract
While talking head generation has advanced rapidly, generating natural listener behavior in dyadic conversations, which know when to react, how to react, and with what type of response, remains underexplored. Existing dyadic datasets lack fine-grained listener reaction annotations, and prevailing evaluation metrics inherited from talking-head and video generation measure visual realism rather than whether a listener reacted appropriately. We address these gaps along three aspects. First, we curate a listening-head-specific dataset built from RealTalk and Seamless Interaction, comprising approximately 147 hours of paired speaker-listener videos with 64,557 event-level reaction annotations across six categories: nodding, head shaking, smiling, laughing, frowning, and surprised. Second, we introduce an audio-driven baseline built on a flow-matching transformer, namely GLARE, with prosody conditioning derived from Qwen2-Audio and a temporal reaction loss that explicitly supervises frame-wise reactions. Third, we propose a reaction-oriented evaluation protocol that jointly measures reaction occurrence (R-F1), temporal alignment (R-tIoU), asymmetric temporal deviation (R-ATD), and reaction-region visual quality (R-FID), giving a more behaviorally grounded assessment than visual-quality-only metrics. Experiment results show consistent gains over prior listening-head methods in both visual fidelity and reaction-level metrics, suggesting that reaction-aware data, modeling, and evaluation are critical for natural listening behavior.