LongEmo: Towards Emotion Understanding and Reasoning in Long Videos

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the bottleneck of multimodal large language models in lacking cumulative dynamic modeling capabilities for long-video affective reasoning by proposing a memory-augmented agent framework. This work introduces LongEmoBench, the first progressive evaluation benchmark tailored for long-video emotion understanding, and designs an event-centric memory architecture. Specifically, it fuses graph neural networks with retrieval-augmented generation to construct an event memory graph, enabling cross-segment long-range dependency modeling and emotion dynamics tracking through the iterative integration of multimodal memories. Experimental results demonstrate that the proposed method achieves state-of-the-art performance against seventeen representative baselines, significantly enhancing long-video affective reasoning capabilities and validating the effectiveness of the proposed architecture.
πŸ“ Abstract
While recent Multimodal Large Language Models (MLLMs) have shown promise in affective computing, their reasoning capabilities are largely confined to short video clips with limited interactions. However, real-world emotions are not merely isolated instantaneous reactions but dynamic and cumulative processes deeply shaped by past experiences and ongoing events. To bridge this gap, we introduce LongEmoBench, a benchmark dedicated to emotion understanding and reasoning in long videos. It assesses progressive capabilities scaling from continuous scene interactions to complex episodic developments. Furthermore, we propose LongEmo, a novel memory-augmented agentic framework designed to tackle the immense challenges of long-range affective reasoning. LongEmo processes continuous video streams to construct an Event Memory Graph, explicitly modeling long-range dependencies and capturing emotional dynamics across discrete events. Given a question, the agent retrieves a query-relevant event stream from the graph, iteratively integrating multimodal memories and relational dependencies to deduce the final answer. Extensive evaluations of 17 representative methods reveal that they struggle significantly with emotion understanding and reasoning in long videos. In contrast, LongEmo achieves state-of-the-art performance, demonstrating the efficacy of its event-centric memory architecture.
Problem

Research questions and friction points this paper is trying to address.

emotion understanding
emotion reasoning
long videos
multimodal large language models
affective computing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Long Video Emotion Reasoning
Memory-Augmented Agentic Framework
Event Memory Graph
Multimodal Large Language Models
Affective Computing
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
S
Shuo Zhang
BUPT
Y
Yifan Zhou
SJTU
H
Han Wang
THU
Jinsong Zhang
Jinsong Zhang
UniversitΓ© Laval
Computer VisionDeep LearningComputer Graphics
J
Jingyu Li
USTC
H
Hongbing Li
BUPT
Zhejun Zhang
Zhejun Zhang
Huawei Zurich Research Center
Autonomous DrivingRoboticsImitation LearningReinforcement Learning
C
Chengyi Zhao
BNU
Y
Yuquan Hao
BUPT
Y
Yitong Liu
BUPT
J
Jiyin Li
BUPT
R
Ruiqi Tang
BUPT
Z
Zixuan Lin
BUPT
Y
Yi Luo
BUPT
X
Xurui Zhang
CUFE
R
Ronghao Chen
PKU
H
Huacan Wang
UCAS
L
Lei Li
BUPT