π€ AI Summary
This study addresses the bottleneck of multimodal large language models in lacking cumulative dynamic modeling capabilities for long-video affective reasoning by proposing a memory-augmented agent framework. This work introduces LongEmoBench, the first progressive evaluation benchmark tailored for long-video emotion understanding, and designs an event-centric memory architecture. Specifically, it fuses graph neural networks with retrieval-augmented generation to construct an event memory graph, enabling cross-segment long-range dependency modeling and emotion dynamics tracking through the iterative integration of multimodal memories. Experimental results demonstrate that the proposed method achieves state-of-the-art performance against seventeen representative baselines, significantly enhancing long-video affective reasoning capabilities and validating the effectiveness of the proposed architecture.
π Abstract
While recent Multimodal Large Language Models (MLLMs) have shown promise in affective computing, their reasoning capabilities are largely confined to short video clips with limited interactions. However, real-world emotions are not merely isolated instantaneous reactions but dynamic and cumulative processes deeply shaped by past experiences and ongoing events. To bridge this gap, we introduce LongEmoBench, a benchmark dedicated to emotion understanding and reasoning in long videos. It assesses progressive capabilities scaling from continuous scene interactions to complex episodic developments. Furthermore, we propose LongEmo, a novel memory-augmented agentic framework designed to tackle the immense challenges of long-range affective reasoning. LongEmo processes continuous video streams to construct an Event Memory Graph, explicitly modeling long-range dependencies and capturing emotional dynamics across discrete events. Given a question, the agent retrieves a query-relevant event stream from the graph, iteratively integrating multimodal memories and relational dependencies to deduce the final answer. Extensive evaluations of 17 representative methods reveal that they struggle significantly with emotion understanding and reasoning in long videos. In contrast, LongEmo achieves state-of-the-art performance, demonstrating the efficacy of its event-centric memory architecture.