VISTA: Value-Informed Event Appraisal for Multimodal Emotion Conflict

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the arbitration challenge in multimodal emotion recognition, where unimodal cues are individually valid yet mutually conflicting. We propose VISTA, a framework built upon Qwen2.5-Omni-7B as the backbone network, which innovatively introduces a seven-field scene-specific appraisal as an intermediate representation to decouple emotional expectations from cue diagnosticity. By integrating log-odds decomposition with joint evidence residual learning, VISTA enables value-based dynamic evidence interpretation and arbitration. Evaluated on the CA-MER dataset, the proposed method improves conflict accuracy to 64.5%, significantly outperforming existing baselines and validating the effectiveness of both the appraisal readout mechanism and downstream decision-making.
📝 Abstract
Conflicting emotional cues can be individually valid: a subdued voice may reflect a blocked goal while a smile satisfies a social obligation. Their interpretation depends on what the event means to the person. We introduce VISTA (Value-Informed Semantic Trust Arbitration), a learned seven-field appraisal interface that conditions modality arbitration on concerns, event relations, and expression conditions while retaining a joint-evidence residual. A log-odds decomposition separates emotion expectation from cue diagnosticity, motivating an interface that lets appraisal change how evidence is interpreted. With a shared Qwen2.5-Omni-7B backbone and matched training examples and steps, VISTA reaches 64.5% conflict accuracy on CA-MER, improving on modality gating by 2.5 percentage points on conflict and 0.2 on consistency. Shuffling appraisal across scenes or removing its decision connection reduces this benefit. A common frozen-backbone probe reaches 0.600 macro CCC for appraisal readout, compared with 0.505 for emotion-only fine-tuning. Evaluations across five benchmarks connect recognition under increasing conflict with appraisal readout and downstream decision use. Together, the analyses and experiments support scene-specific appraisal as an intermediate representation that helps interpret conflicting emotional evidence.
Problem

Research questions and friction points this paper is trying to address.

multimodal emotion conflict
emotion recognition
conflicting emotional cues
appraisal
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Emotion Conflict
Appraisal Interface
Log-odds Decomposition
Modality Arbitration
Intermediate Representation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jiale Dai
State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University
L
Liuxian Ma
College of Artificial Intelligence, Tsinghua University
X
Xiaoke Niu
State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University
W
Wenjing Zhang
China United Network Communications Group Co., Ltd.
H
Huiying Zhao
China United Network Communications Group Co., Ltd.
Zhaoxiang Liu
Zhaoxiang Liu
China Unicom
Computer VisionDeep LearningRoboticsHuman-Computer Interaction
Shiguo Lian
Shiguo Lian
CloudMinds
Guojie Song
Guojie Song
Professor (Research), Tenured of Peking University
Psychological AIAI Safe & Value AlignmentAgent Cognition & Behavioral ModelingLLM&GML