Seek-and-View Reasoning for Multi-View Spatial Understanding

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the fragility of cross-view alignment and the geometry-to-language translation bottleneck in existing multi-view spatial reasoning by proposing Vantage, a novel framework that introduces a "Seek-and-View" paradigm. Through dual viewpoint and geometric guidance, Vantage pairs vision-language models (VLMs) with 3D foundation models to sequentially perform viewpoint-guided reasoning and geometry-guided evidence synthesis, thereby precisely localizing implicit spatial evidence. The proposed framework operates in a plug-and-play manner without requiring fine-tuning. Extensive evaluations demonstrate consistent performance improvements across six VLMs and five benchmarks. By significantly reducing reliance on language alignment, Vantage effectively enhances multi-view spatial understanding capabilities.
📝 Abstract
Existing approaches to multi-view spatial reasoning operate largely on sparse input views. Vision-language models (VLMs) are thus restricted to understand a scene and infer spatial relations within these fixed views, leading to fragile cross-view alignment and geometry-to-language bottleneck. To address these issues, we formulate a novel Seek-and-View reasoning approach to find implicit cross-view spatial evidence by locating a question-relevant view to support the spatial reasoning. To realize this approach, we propose Vantage, a training-free model-agnostic reasoning framework that pairs a VLM with a 3D foundation model: a viewpoint-grounded reasoning stage for question analysis and view planning, followed by a geometry-grounded evidence augmentation stage to effectively synthesize and incorporate visual evidence into the final reasoning. Comprehensive experiments on six VLMs demonstrate consistent improvements on five benchmarks without fine-tuning. Overall, by revealing spatial evidence through view-grounded reasoning, Vantage can largely reduce reliance on language-based cross-view alignment and improve multi-view spatial understanding. Our code is available at https://github.com/q1xiangchen/Vantage.
Problem

Research questions and friction points this paper is trying to address.

multi-view spatial reasoning
vision-language models
cross-view alignment
sparse input views
Innovation

Methods, ideas, or system contributions that make the work stand out.

Seek-and-View Reasoning
Multi-View Spatial Understanding
Training-free Framework
Vision-Language Models
3D Foundation Model
🔎 Similar Papers
No similar papers found.