🤖 AI Summary
This study addresses the vulnerability of multimodal emotion understanding to the “Clever Hans effect,” wherein models rely on superficial cues and exhibit insufficient reliability in implicit or conflicting scenarios. Inspired by cognitive appraisal theory, this work proposes a novel “perception–appraisal” paradigm. It designs an interleaved Mixture-of-Experts (MoE) architecture to enable efficient adaptation with limited parameters and introduces the AEQS metric to quantify the quality of appraisal evidence. The contributions encompass a comprehensive framework integrating datasets, models, and benchmarks. Achieving state-of-the-art performance on CogEmo-Bench, the proposed approach demonstrates strong cross-domain generalization capabilities, effectively advancing multimodal large language models toward deeper emotional reasoning and closer alignment with human cognition.
📝 Abstract
Recent multimodal large language models (MLLMs) increasingly incorporate explainable reasoning for emotion understanding. However, reasoning based mainly on observable affective cues can reduce emotion understanding to superficial cue-label associations, giving rise to the Clever Hans effect. Such shortcuts become unreliable when affective cues are implicit, conflicting across modalities, linguistically misleading, or obscured by redundant details. In contrast, human emotions are shaped by how individuals interpret and evaluate surrounding events beyond observable cues. Inspired by appraisal theories of emotion, we formulate multimodal emotion understanding as a progression from perception to cognitive appraisal, and introduce a dataset, a model, and a benchmark to support this novel paradigm. CogEmo-40K is a large-scale instruction-tuning dataset constructed through a perception-to-appraisal pipeline to elicit evidence-grounded reasoning across six cognitive appraisal dimensions underlying emotion. CogEmo-MoE is a compact sparse MLLM that introduces interleaved MoE blocks for appraisal-specific adaptation, enabling effective appraisal reasoning at a substantially smaller scale than typical emotion MLLMs. CogEmo-Bench introduces an Appraisal Evidence Quality Score (AEQS) to assess cognitive-affective understanding across six complementary appraisal dimensions, addressing the limitation of conventional emotion metrics that evaluate what emotion is predicted but not why it arises. Extensive experiments show that our paradigm not only leads CogEmo-Bench, but also exhibits strong cross-domain generalization. Our findings suggest that perception-to-appraisal reasoning can move beyond surface-level cue-label associations toward more reliable multimodal emotion understanding and closer cognitive alignment between MLLMs and humans.