MultiVENT-Raw: A Benchmark for Retrieval and Reasoning over Raw Videos

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决原始视频信息检索和理解难题,发布包含近12万条视频的多语言数据集MultiVENT-Raw,并支持检索和生成任务。
📝 Abstract
Online information is increasingly consumed in video format. Much of this comes in the form of *raw video*: continuous footage taken on a cell phone, with a hand-held camera, or via CCTV, which is then directly uploaded to social media platforms and content sharing services. Whereas professional or even amateur-edited footage tends to feature scripted speech, chyrons, graphics, and metadata that help contextualize its subject matter, raw video typically contains none of these things, making it a much more challenging medium for information retrieval and machine understanding. To facilitate progress in this domain, we release MultiVENT-Raw, a multilingual collection of nearly 120,000 primarily raw videos (over 5,300 total hours), paired with 130 events and 222 event-centric queries, along with human-annotated video relevance judgments and human-extracted key facts for relevant videos. MultiVENT-Raw supports both a retrieval task---to identify videos in the collection relevant to a query event---and a generation task---to summarize event-related videos into a coherent report for a target user. We benchmark strong baselines on MultiVENT-Raw, showing both tasks to be challenging even for some of the latest multimodal models.
Problem

Research questions and friction points this paper is trying to address.

raw video
information retrieval
machine understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

raw video
multimodal models
information retrieval
machine understanding
benchmark
🔎 Similar Papers
No similar papers found.
Reno Kriz
Reno Kriz
Associate Research Scientist
information extractionvideo retrievaltext simplificationlarge language models
David Etter
David Etter
Machine Learning Scientist
Machine learning / Deep LearningComputer VisionNatural Language Processing (NLP)and Information Retrieval
Alexander Martin
Alexander Martin
Johns Hopkins University
Multimodal RAGMultimodal ReasoningVideo Understanding
C
Cameron Carpenter
Johns Hopkins University
D
Debashish Chakraborty
Human Language Technology Center of Excellence, Johns Hopkins University
H
Hannah Recknor
Human Language Technology Center of Excellence, Johns Hopkins University
Reihaneh Iranmanesh
Reihaneh Iranmanesh
Computer Science and Statistics, Amherst College
Human-Robot InteractionMLAI AlignmentMechanistic Interpretability‬
Matthew Maciejewski
Matthew Maciejewski
Johns Hopkins University
speech separationspeaker diarizationspeaker identification
Kenton Murray
Kenton Murray
Research Scientist, Johns Hopkins
Machine LearningNatural Language ProcessingMachine TranslationSemanticsNeural Networks
Eugene Yang
Eugene Yang
Research Scientist, Johns Hopkins University, Human Language Technology Center of Excellence
High Recall RetrievalCross-lingual Information RetrievalInformation RetrievaleDiscovery
Benjamin Van Durme
Benjamin Van Durme
Johns Hopkins University / Microsoft
LinguisticsNatural Language ProcessingArtificial Intelligence
Aaron Steven White
Aaron Steven White
University of Rochester
Andrew Yates
Andrew Yates
Johns Hopkins University, Human Language Technology Center of Excellence
Information RetrievalNLPAI
William Walden
William Walden
Research Scientist, Johns Hopkins University HLTCOE
Natural Language ProcessingEvaluationAI for Science