MVVBench: Benchmarking 4D Reasoning in Vision-Language Models

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of joint cross-view and temporal reasoning in vision-language models (VLMs) applied to multi-view videos. We construct an evaluation benchmark based on real-world multi-camera data, designing single-view-insolvable tasks to assess six core capabilities and revealing inherent deficiencies in existing models regarding identity tracking and multi-hop reasoning. Methodologically, we propose an evaluation paradigm that enforces joint reasoning, introducing chain-of-thought scaffolding, structured cross-view evidence aggregation, and reinforcement learning with verifiable rewards. Experimental results demonstrate that these inference-time strategies significantly enhance performance without retraining. Furthermore, our findings provide preliminary validation that reinforcement learning can elicit latent multi-view capabilities, thereby laying a foundation for reliable embodied perception.
📝 Abstract
Multi-view video understanding requires integrating spatial and temporal evidence across multiple, often non-overlapping camera streams: tracking entities as they transition between viewpoints, aligning events across time, and reasoning about latent 4D continuity rather than any single visible frame. We introduce MVVBench, a benchmark for multi-view video reasoning built from real world multi camera datasets. Questions are curated to be monocular-ambiguous along both the view and the temporal axis: each question is unanswerable from any single view in the designated input set, and the majority are further unanswerable from any single moment. Each question becomes uniquely solvable only by jointly reasoning across views and across time. MVVBench spans diverse dynamic scenes and probes six capabilities: implicit/explicit attribute identification, implicit/explicit relative distance, relative camera pose, and compositional counting, with human-authored QA and rigorous verification. Beyond benchmarking, we provide an extensive analysis of when and why current vision language models succeed or fail, characterizing errors due to temporal mis-localization, cross-view identity breaks, and brittle multi-hop reasoning. We then study inference-time elicitation strategies that unlock latent multi-view competence---task-specific chain-of-thought scaffolds and structured cross-view evidence aggregation---yielding substantial gains without retraining. Finally, we present preliminary evidence that reinforcement learning with verifiable rewards can elicit some latent multi-view competence in the base model, pointing to training-time approaches as a promising direction for future work. Together, MVVBench offers a rigorous evaluation of 4D multi-view reasoning and a foundation for future progress toward reliable embodied perception.
Problem

Research questions and friction points this paper is trying to address.

multi-view video understanding
4D reasoning
vision-language models
spatiotemporal integration
benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-view Video Reasoning
4D Spatiotemporal Benchmark
Chain-of-Thought Scaffolding
Cross-view Evidence Aggregation
Reinforcement Learning