ReVA: A Scene-Centric Dataset Beyond Repetition for Remote Sensing Video Question Answering

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the limitations of existing remote sensing benchmarks, which rely on static templates and overlook video temporal dynamics, thereby inadequately evaluating spatiotemporal reasoning capabilities. To bridge this gap, we construct a UAV video question-answering dataset comprising 2,438 videos across 18 global cities and 22,000 high-quality QA pairs, annotated via a semi-automatic pipeline integrating large language models and multimodal models to ensure data quality. Furthermore, this study introduces the first scene-centric evaluation framework for remote sensing videos, transcending static image constraints and filling the critical void in temporal reasoning assessment. Through systematic benchmarking of 23 mainstream video foundation models, we reveal current technological bottlenecks and establish a pivotal benchmark for advancing remote sensing video understanding.
πŸ“ Abstract
Multimodal Large Language Models (MLLMs) have demonstrated remarkable advances in remote sensing. However, existing remote sensing multimodal reasoning benchmarks exhibit two critical limitations: they rely on (i) template-driven questions, which causes repetitive questions; and (ii) static images that fail to capture the inherent temporal nature of drone/UAV videos. This leaves systematic evaluation of remote sensing video reasoning largely unexplored. To address this gap, we introduce ReVA, a new dataset for remote sensing video question answering, designed to assess spatiotemporal, scene-centric, and reasoning-oriented capabilities of MLLMs. ReVA comprises 2,438 drone videos spanning 18 cities worldwide (580K frames) and 22K high-quality question-answer pairs across 11 challenging QA tasks. We develop a semi-automatic annotation pipeline that leverages Text LLMs and MLLMs for question-answer generation with human verification. We evaluate 23 proprietary and open-source Video LLMs on ReVA, exposing fundamental limitations of current models. These findings position ReVA as a critical benchmark toward better remote sensing video understanding and temporal reasoning capabilities for real-world deployments. Our code and dataset are available at: https://github.com/zyaocoder/ReVA
Problem

Research questions and friction points this paper is trying to address.

Remote Sensing
Video Question Answering
Multimodal Large Language Models
Temporal Reasoning
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Remote Sensing Video Question Answering
Multimodal Large Language Models
Semi-automatic Annotation Pipeline
Spatiotemporal Reasoning
Scene-Centric Dataset
πŸ”Ž Similar Papers
No similar papers found.