🤖 AI Summary
This study addresses visual hallucination and the neglect of subjectivity in multimodal large language models for affective reasoning, which arise from hard-label supervision. To this end, we propose a Valence-Arousal-Dominance (VAD) space-based affective distribution prior framework. Methodologically, we design a distribution-aligned affective diversity reward mechanism to enhance subjective emotion coverage and introduce a counterfactual visual intervention gating technique to suppress interpretations lacking visual evidence, thereby jointly optimizing affective diversity and visual alignment. By integrating reinforcement learning with probabilistic distribution matching, our approach achieves state-of-the-art performance across multiple public benchmarks, improving cross-domain average accuracy by 10.8% over EMO-R3.
📝 Abstract
Reinforcement Learning has significantly advanced the complex reasoning capabilities of MLLMs. However, prevailing RL algorithms suffer a severe failure in emotion reasoning tasks. These methods heavily rely on deterministic hard-label supervision and point-wise isolated evaluation, creating a fundamental gap with the inherently subjective and continuously distributed nature of human emotions. Furthermore, unlike explicit physical objects, emotional states are deeply implicit within visual cues. This abstract nature exacerbates visual hallucinations in MLLMs, leading to plausible yet ungrounded emotional evidence. To address these limitations, we propose Diversity-Aware Subjective Policy Optimization (DSPO), a reinforcement learning framework that jointly promotes subjective affective coverage and visual grounding. First, we construct a context-grounded emotional distribution prior in the VAD space by combining the lexical prior of the annotated emotion with image-specific contextual information. Based on this prior, we introduce a Distribution-Aligned Emotional Diversity Reward (DEDR), which measures the leave-one-out marginal contribution of each candidate emotion within a rollout. DEDR rewards candidates whose inclusion brings the predicted affective set closer to the context-grounded prior, thereby preserving plausible subjective interpretations without encouraging unconstrained dispersion. We further develop Counterfactual Visual Intervention Gating (CVIG), which masks the visual region highlighted in the reasoning process and uses the resulting candidate-wise probability changes to reduce the weights of interpretations unsupported by visual evidence. Extensive experiments demonstrate that DSPO achieves state-of-the-art performance across multiple public benchmarks, especially on the cross-domain performance, i.e., improving +10.8\% on average cross-domain accuracy than EMO-R3.