🤖 AI Summary
This study addresses the inability of visual question answering (VQA) models to reliably abstain when evidence is insufficient, which frequently leads to erroneous outputs. We construct a unified benchmark spanning chart and scene domains to jointly evaluate correct answering and necessary abstention, ensuring strict alignment between answers and supporting evidence. Furthermore, we propose a holistic task success rate metric, revealing that evaluating abstention decisions in isolation obscures underlying answering errors. Supervision labels are generated using executable witnesses and residual cue analysis based on the PlotQA, CLEVR, and GQA datasets. Experimental results demonstrate that the best-performing model achieves a holistic task success rate of only 57%; notably, even with an abstention accuracy of 96.2%, numerous sample groups containing incorrect answers persist.
📝 Abstract
Reliable visual question answering requires correct answers when evidence is sufficient and abstention when it is not. We introduce a benchmark that connects complete-question evaluation with explicit evidence for its labels across PlotQA charts, CLEVR rendered scenes, and GQA photographs. Each question groups original and edited images, presented independently; success requires every supported answer and every required abstention to be correct. For chart missing-information labels, executable witnesses establish that admissible complete charts give different answers but identical pixels after masking. Scene labels follow source programs and edits, with a residual-cue analysis for photographs. Across 72,000 responses from six model configurations, the highest observed complete task success rates are 57.0%, 43.5%, and 33.7%, respectively. On charts, the strongest configuration achieves 96.2% per-view decision accuracy, yet 265 of its 835 groups with every decision correct still contain incorrect answers. Evaluating supported answers and necessary abstentions together exposes failures that answerability decisions alone conceal.