๐ค AI Summary
This work addresses a critical limitation in existing multimodal conversational emotion recognition approachesโtheir neglect of emotional inertia, which hinders performance gains by overlooking the influence of prior emotional states on current emotion dynamics. To tackle this issue, the study introduces emotional inertia modeling into the task for the first time and proposes an Emotional Inertia-Informed Supervised Contrastive Learning (EII-SCL) module. This module constructs positive and negative samples via temporal window sampling to reflect inertia effects, explicitly incorporates emotional inertia as a prior into the contrastive learning objective, and seamlessly integrates into existing architectures without requiring additional data. Experimental results on the IEMOCAP and MELD datasets demonstrate that the proposed method significantly outperforms state-of-the-art approaches, achieving notable improvements in emotion recognition accuracy.
๐ Abstract
Multimodal emotion recognition in conversation (MERC) achieves accurate predictions by integrating multimodal and contextual information in dialogues. While current MERC approaches focus on modeling complex contextual dependencies in conversation, they often overlook the impact of contextual emotional inertia in emotion shift, leading to sub-optimal performance. To address this issue, we propose a novel Emotional Inertia-Informed Supervised Contrastive Learning module (EII-SCL) that informs the contrastive objective by constructing inertia-affected samples within temporal windows, effectively leveraging emotional inertia as a prior while enabling seamless integration with existing MERC models without requiring additional data. Extensive experiments on IEMOCAP and MELD show that our approach consistently outperforms state-of-the-art methods.