From Expression to Reaction: Role-aware Visual Transfer and Stimulus-guided Reasoning for Interlocutor Emotion Recognition

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of annotated data for listener emotion recognition in conversations and the ambiguity arising from visually similar reactions. To tackle these challenges, this work proposes the RASG framework, which trains a listener-expert model through role-aware visual transfer and a novel pseudo-labeling strategy. Furthermore, it introduces a stimulus-guided linguistic reasoning mechanism that is selectively activated only under uncertainty to resolve such ambiguities. Evaluated on the MER-Cross dataset, the proposed framework achieves an accuracy of 76.25%, yielding a performance improvement of over 17% and securing second place in the ACM MM 2026 Challenge.
📝 Abstract
In this paper, we propose a Role-aware Stimulus-guided (RASG) framework for interlocutor emotion recognition, which predicts listener emotions from listener-only videos and speaker-only audios. RASG consists of Role-aware Visual Transfer (RVT) and Stimulus-guided Boundary Reasoning (SBR) modules, which address supervision mismatch due to the lack of labeled listener data and ambiguity among visually similar listener reactions whose interpretation depends on speaker context, respectively. More specifically, RVT selects speaker samples whose facial expressions support their emotion labels. It then filters listener tracks and uses reliable pseudo-labels to train a listener-centric visual expert. SBR uses a two-class language reasoner only when the visual model is uncertain. It treats speaker audio and text as context rather than direct emotion evidence to distinguish similar listener reactions. Experiments conducted on MER-Cross dataset shows that RASG achieves 76.25\% on MER-Cross and improves the performance of the baseline over 17\%. Our team ranks second in Track 1 (MER-Cross) of the MER Grand Challenge at ACM MM 2026.
Problem

Research questions and friction points this paper is trying to address.

Interlocutor Emotion Recognition
Supervision Mismatch
Visual Ambiguity
Listener Reaction
Multimodal Emotion Recognition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Interlocutor Emotion Recognition
Role-aware Visual Transfer
Stimulus-guided Boundary Reasoning
Pseudo-labels
Multimodal Reasoning
🔎 Similar Papers
2024-05-14IEEE/RJS International Conference on Intelligent RObots and SystemsCitations: 2
W
Wei Wang
Guangdong University of Technology
Z
Zhaowu Li
Guangdong University of Technology
J
Jianjie Luo
Guangdong University of Technology
Fu Lee Wang
Fu Lee Wang
Hong Kong Metropolitan University
AIData ScienceLearning Technology
Lap-Kei Lee
Lap-Kei Lee
Hong Kong Metropolitan University
Z
Zhenguo Yang
Guangdong University of Technology