When Does an Image Determine the Answer? Benchmarking Visual Answerability across Charts and Scenes

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inability of visual question answering (VQA) models to reliably abstain when evidence is insufficient, which frequently leads to erroneous outputs. We construct a unified benchmark spanning chart and scene domains to jointly evaluate correct answering and necessary abstention, ensuring strict alignment between answers and supporting evidence. Furthermore, we propose a holistic task success rate metric, revealing that evaluating abstention decisions in isolation obscures underlying answering errors. Supervision labels are generated using executable witnesses and residual cue analysis based on the PlotQA, CLEVR, and GQA datasets. Experimental results demonstrate that the best-performing model achieves a holistic task success rate of only 57%; notably, even with an abstention accuracy of 96.2%, numerous sample groups containing incorrect answers persist.
📝 Abstract
Reliable visual question answering requires correct answers when evidence is sufficient and abstention when it is not. We introduce a benchmark that connects complete-question evaluation with explicit evidence for its labels across PlotQA charts, CLEVR rendered scenes, and GQA photographs. Each question groups original and edited images, presented independently; success requires every supported answer and every required abstention to be correct. For chart missing-information labels, executable witnesses establish that admissible complete charts give different answers but identical pixels after masking. Scene labels follow source programs and edits, with a residual-cue analysis for photographs. Across 72,000 responses from six model configurations, the highest observed complete task success rates are 57.0%, 43.5%, and 33.7%, respectively. On charts, the strongest configuration achieves 96.2% per-view decision accuracy, yet 265 of its 835 groups with every decision correct still contain incorrect answers. Evaluating supported answers and necessary abstentions together exposes failures that answerability decisions alone conceal.
Problem

Research questions and friction points this paper is trying to address.

Visual Question Answering
Visual Answerability
Abstention
Benchmark
Evidence Sufficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Answerability
Benchmark
Executable Witnesses
Abstention
Residual-cue Analysis
🔎 Similar Papers
S
Sungguk Cha
LG Uplus, Seoul, South Korea
M
Mintae Kim
LG Uplus, Seoul, South Korea
Y
Youngsub Han
LG Uplus, Seoul, South Korea
B
Byoung-Ki Jeon
LG Uplus, Seoul, South Korea
Sangyeob Lee
Sangyeob Lee
Associate Professor, Dept. of Materials Science and Engineering, Hanbat National University
Materials ScienceOLED2D MaterialsGrapheneKelvin Probe Force Microscopy