🤖 AI Summary
This study addresses the bottleneck wherein multimodal large language models struggle to distinguish semantically similar emotions based on fine-grained visual evidence. To this end, we propose DAN, a training-free inference optimization framework that introduces a novel test-time optimization paradigm. Specifically, DAN employs a Hierarchical Emotion Reasoning Chain (HERC) to guide fine-grained emotion attribution and integrates Contrastive Discriminative Visual Pruning (CDVP) to eliminate redundant information, thereby enhancing discriminative capability. This design effectively decouples and synergistically improves the model's attribution and discrimination performance. Experimental results demonstrate that DAN significantly enhances recognition accuracy for subtle emotional nuances, achieving a 10.47% performance gain on the WebEmo25 benchmark.
📝 Abstract
While multimodal large language models (MLLMs) have demonstrated exceptional capabilities in objective understanding tasks, their performance in affective reasoning still falls significantly short of human standards. We attribute it to a central capability gap: MLLMs are difficult to reliably distinguish semantically proximal emotions based on fine-grained visual evidence, which could be decoupled as two limitations: 1) Insufficient Attribution. The global reasoning paradigm of conventional MLLMs severely dilutes fine-grained emotion cues, where subtle emotional states are usually implicitly encoded, thereby generating emotional misjudgments in complex scenarios. 2) Insufficient Discrimination. Existing methods could only identify regions generally associated with emotions, which fails to distinguish discriminative regions between semantically similar emotions, leading to ambiguous emotion judgements. To overcome these limitations, we present a training-free inference-time optimization framework, named Decoding Affective Nuances (DAN). Specifically, we propose a Hierarchical Emotional Reasoning Chain (HERC) that enhances the insufficient attribution by harmonizing fine-grained scene/object-level cues and performing a soft-gated reasoning. Furthermore, to discriminate between semantically proximal emotions, we design a Contrastive Discriminative Visual Pruning (CDVP), which isolates discriminative visual tokens to reason the final emotion category by computing the absolute discrepancy between the attention distributions of similar emotions. Performances on several benchmarks demonstrate that DAN significantly improves discrimination for affective nuances without consuming additional training resources, especially achieving +10.47% improvements with Qwen3-VL-8B-Instruct on WebEmo25 dataset that contains 25 fine-grained emotion categories.