🤖 AI Summary
This study addresses the scarcity of annotated data for listener emotion recognition in conversations and the ambiguity arising from visually similar reactions. To tackle these challenges, this work proposes the RASG framework, which trains a listener-expert model through role-aware visual transfer and a novel pseudo-labeling strategy. Furthermore, it introduces a stimulus-guided linguistic reasoning mechanism that is selectively activated only under uncertainty to resolve such ambiguities. Evaluated on the MER-Cross dataset, the proposed framework achieves an accuracy of 76.25%, yielding a performance improvement of over 17% and securing second place in the ACM MM 2026 Challenge.
📝 Abstract
In this paper, we propose a Role-aware Stimulus-guided (RASG) framework for interlocutor emotion recognition, which predicts listener emotions from listener-only videos and speaker-only audios. RASG consists of Role-aware Visual Transfer (RVT) and Stimulus-guided Boundary Reasoning (SBR) modules, which address supervision mismatch due to the lack of labeled listener data and ambiguity among visually similar listener reactions whose interpretation depends on speaker context, respectively. More specifically, RVT selects speaker samples whose facial expressions support their emotion labels. It then filters listener tracks and uses reliable pseudo-labels to train a listener-centric visual expert. SBR uses a two-class language reasoner only when the visual model is uncertain. It treats speaker audio and text as context rather than direct emotion evidence to distinguish similar listener reactions. Experiments conducted on MER-Cross dataset shows that RASG achieves 76.25\% on MER-Cross and improves the performance of the baseline over 17\%. Our team ranks second in Track 1 (MER-Cross) of the MER Grand Challenge at ACM MM 2026.