GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the dual challenges of answer accuracy and evidence view localization in multi-view driving scenario visual question answering by proposing CoVeR-VQA, a training-free framework. The method integrates multiple multimodal large language models, including GPT and Gemini, for zero-shot prediction. Furthermore, it introduces a novel cross-segment group-level verification mechanism that performs multi-stage stepwise correction through semantic filtering and prior knowledge integration, thereby enabling precise reasoning and evidence localization. Experimental results demonstrate that the proposed framework achieves a joint accuracy of 84.75%, outperforming the baseline by 13.56 percentage points, with the final submission reaching an accuracy of 88.14%.
📝 Abstract
GoldenViewVQA requires models to jointly answer driving-scene questions and identify the camera view containing the supporting visual evidence, making precise evidence localization as important as answer correctness. We present \textbf{CoVeR-VQA}, a training-free multi-stage verification and correction framework for grounded multi-view VQA. Starting from GPT-5.6 zero-shot predictions, CoVeR-VQA progressively applies view-specific verification with Gemini-3.6-Flash, prior-guided joint verification with Claude-Opus-5, and cross-split group-level verification that exploits semantically filtered question groups from shared multi-view scenes and validation-derived prior knowledge. On the official GoldenViewVQA test set, the four-stage CoVeR-VQA pipeline achieves 84.75\% Joint Accuracy, improving the GPT-5.6 zero-shot baseline by 13.56 percentage points, while reaching 94.92\% Answer Accuracy and 86.44\% View Accuracy. The final submitted run achieves 88.14\% Joint Accuracy after two additional evaluator-informed post-hoc corrections. Our analysis shows that supporting-view localization remains the primary source of residual errors, highlighting the importance of explicit evidence verification for reliable multi-view multimodal reasoning.
Problem

Research questions and friction points this paper is trying to address.

Multi-view VQA
Evidence Localization
Driving Scene Understanding
Grounded Visual Question Answering
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-free Framework
Multi-stage Verification
Multi-view VQA
Evidence Localization
LLM Ensemble
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kun Wang
School of Software, Shandong University, Jinan, China
Yupeng Hu
Yupeng Hu
Shandong University
Multimedia Information RetrievalData Mining and Knowledge Discovery
R
Ruping Cao
School of Software, Shandong University, Jinan, China
Hao Liu
Hao Liu
University of Electronic Science and Technology of China
RISstacked intelligent metasurfaceDRL
Z
Zhiran Li
School of Software, Shandong University, Jinan, China
Q
Qianlong Xiang
School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen), Shenzhen, China
Harry Cheng
Harry Cheng
National University of Singapore
Diffusion ModelMLLM SecurityDeepfake Detection