VepAgent: Bridging Causal-Transition via Tool-Augmented Reinforcement Learning for Video Event Prediction

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of multimodal large language models in predicting unobserved causal transitions in videos due to their reliance on retrospective summarization. We propose an agent framework integrating causal reasoning with tool-augmented reinforcement learning to explicitly model the logical progression from current states to future events. Key contributions include constructing the FutureBench-4K dataset to bridge gaps in causal logic, designing a dynamic diagnostic toolkit for spatiotemporal evidence recovery, and employing a composite reward mechanism to optimize causal consistency. By incorporating supervised fine-tuning, state tracking, and multimodal visual grounding, the proposed framework achieves state-of-the-art performance on both FutureBench and NEVBench, significantly outperforming larger-scale multimodal models.
📝 Abstract
Multimodal Large Language Models (MLLMs) have demonstrated remarkable potential in video understanding, yet their reliance on retrospective summarization and text-centric priors often limits their ability to bridge unobserved causal transitions when applied to Video Event Prediction (VEP). To address this, we propose VepAgent, an agentic framework that integrates causal-transition reasoning with tool-augmented reinforcement learning (RL) for robust VEP. Unlike prior methods that passively project future trajectories from historical dependencies, our approach explicitly models the logical progression from terminal observed states to future events. Specifically, we first construct futurebench-4K, a high-quality chain-of-thought dataset for supervised fine-tuning (SFT) that effectively bridges the causal-logic gap by structuring the deduction of unobserved intermediate states. Subsequently, we develop a diagnostic tool library integrating state tracking, frame retrieval, and region magnification, enabling the agent to dynamically augment reasoning with external tools to recover missing spatio-temporal evidence and resolve visual ambiguities during inference. Moreover, we propose a composite reward mechanism that jointly optimizes prediction accuracy, causal coherence, and reliable prior, compelling the agent to rely on genuine visual grounding rather than superficial textual similarities. Extensive evaluations on FutureBench and NEPBench datasets demonstrate that our method achieves state-of-the-art performance, significantly outperforming larger MLLMs and validating the empirical effectiveness of our agentic, future-oriented reasoning paradigm.
Problem

Research questions and friction points this paper is trying to address.

Video Event Prediction
Multimodal Large Language Models
Causal-Transition Reasoning
Unobserved States
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video Event Prediction
Tool-Augmented Reinforcement Learning
Causal-Transition Reasoning
Chain-of-Thought
Multimodal Large Language Models
🔎 Similar Papers
Q
Qiutong Chen
Nankai University
Y
Yuchan Guo
Carnegie Mellon University
Z
Zhenlong Yuan
Xiaohongshu Inc.
H
Haobo Yang
Columbia University
F
Fangfang Lin
Santa Clara University
X
Xinyi Long
Xiaohongshu Inc.
Y
Yin Wang
New York University
Z
Zijian Song
Xiaohongshu Inc.
R
Rui Lan
Xiaohongshu Inc.
S
Shi Qiu
Xiaohongshu Inc.
Boyuan Pan
Boyuan Pan
TechLead, RedNote Inc.
Natural Language ProcessingSearch EngineRecommendation Systems
Y
Yang Luo
Xiaohongshu Inc.
Yuyin Zhou
Yuyin Zhou
Assistant Professor, Computer Science and Engineering, Genomics Institute, UC Santa Cruz
medical image analysismachine learningcomputer visionAI in healthcare