DSPO: Diversity-aware Subjective Policy Optimization for Robust Emotional Reasoning

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses visual hallucination and the neglect of subjectivity in multimodal large language models for affective reasoning, which arise from hard-label supervision. To this end, we propose a Valence-Arousal-Dominance (VAD) space-based affective distribution prior framework. Methodologically, we design a distribution-aligned affective diversity reward mechanism to enhance subjective emotion coverage and introduce a counterfactual visual intervention gating technique to suppress interpretations lacking visual evidence, thereby jointly optimizing affective diversity and visual alignment. By integrating reinforcement learning with probabilistic distribution matching, our approach achieves state-of-the-art performance across multiple public benchmarks, improving cross-domain average accuracy by 10.8% over EMO-R3.
📝 Abstract
Reinforcement Learning has significantly advanced the complex reasoning capabilities of MLLMs. However, prevailing RL algorithms suffer a severe failure in emotion reasoning tasks. These methods heavily rely on deterministic hard-label supervision and point-wise isolated evaluation, creating a fundamental gap with the inherently subjective and continuously distributed nature of human emotions. Furthermore, unlike explicit physical objects, emotional states are deeply implicit within visual cues. This abstract nature exacerbates visual hallucinations in MLLMs, leading to plausible yet ungrounded emotional evidence. To address these limitations, we propose Diversity-Aware Subjective Policy Optimization (DSPO), a reinforcement learning framework that jointly promotes subjective affective coverage and visual grounding. First, we construct a context-grounded emotional distribution prior in the VAD space by combining the lexical prior of the annotated emotion with image-specific contextual information. Based on this prior, we introduce a Distribution-Aligned Emotional Diversity Reward (DEDR), which measures the leave-one-out marginal contribution of each candidate emotion within a rollout. DEDR rewards candidates whose inclusion brings the predicted affective set closer to the context-grounded prior, thereby preserving plausible subjective interpretations without encouraging unconstrained dispersion. We further develop Counterfactual Visual Intervention Gating (CVIG), which masks the visual region highlighted in the reasoning process and uses the resulting candidate-wise probability changes to reduce the weights of interpretations unsupported by visual evidence. Extensive experiments demonstrate that DSPO achieves state-of-the-art performance across multiple public benchmarks, especially on the cross-domain performance, i.e., improving +10.8\% on average cross-domain accuracy than EMO-R3.
Problem

Research questions and friction points this paper is trying to address.

Emotional Reasoning
Reinforcement Learning
Multimodal Large Language Models
Visual Hallucination
Subjectivity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Subjective Policy Optimization
Emotional Diversity Reward
Counterfactual Visual Intervention
Visual Grounding
Reinforcement Learning
C
Cheng Ye
University of Science and Technology of China, Hefei
W
Weidong Chen
University of Science and Technology of China, Hefei
B
Bingyan Xu
University of Science and Technology of China, Hefei
Zhendong Mao
Zhendong Mao
University of Science and Technology of China
CV,NLP