VB: Visibility Benchmark for Visibility and Perspective Reasoning in Images

📅 2026-03-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether vision-language models can accurately assess the visibility of image content and proactively abstain when evidence is insufficient. To this end, the authors construct a benchmark comprising 300 rigorously curated evaluation items, leveraging minimally edited image–text pairs to probe models’ judgment and abstention capabilities. They introduce “reason encoding” to explain unanswerability and employ controlled minimal edits to verify whether model responses shift appropriately with changes in evidential support. The work proposes a suite of multidimensional evaluation metrics: Confidence-Aware Accuracy (CAA), Minimal-Edit Flip Rate (MEFR), Selective Prediction via Confidence Ranking (SelRank), and Theory-of-Mind-inspired Accuracy (ToMAcc). Experiments reveal that GPT-4o and Gemini 1.5 Pro achieve the best overall performance (aggregate score ≈ 0.728), while the open-source Gemma 3 12B model (0.505) outperforms several older closed-source systems.

Technology Category

Computer Vision: Language and VisionMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSecurity and Privacy: Large-scale security measurements
📝 Abstract
We present VB, a benchmark that tests whether vision-language models can determine what is and is not visible in a photograph, and abstain when a human viewer cannot reliably answer. Each item pairs a single photo with a short yes/no visibility claim; the model must output VISIBLY_TRUE, VISIBLY_FALSE, or ABSTAIN, together with a confidence score. Items are organized into 100 families using a 2x2 design that crosses a minimal image edit with a minimal text edit, yielding 300 headline evaluation cells. Unlike prior unanswerable-VQA benchmarks, VB tests not only whether a question is unanswerable but why (via reason codes tied to specific visibility factors), and uses controlled minimal edits to verify that model judgments change when and only when the underlying evidence changes. We score models on confidence-aware accuracy with abstention (CAA), minimal-edit flip rate (MEFR), confidence-ranked selective prediction (SelRank), and second-order perspective reasoning (ToMAcc); all headline numbers are computed on the strict XOR subset (three cells per family, 300 scored items per model). We evaluate nine models spanning flagship and prior-generation closed-source systems, and open-source models from 8B to 12B parameters. GPT-4o and Gemini 3.1 Pro effectively tie for the best composite score (0.728 and 0.727), followed by Gemini 2.5 Pro (0.678). The best open-source model, Gemma 3 12B (0.505), surpasses one prior-generation closed-source system. Text-flip robustness exceeds image-flip robustness for six of nine models, and confidence calibration varies substantially: GPT-4o and Gemini 2.5 Pro achieve similar accuracy yet differ sharply in selective prediction quality.
Problem

Research questions and friction points this paper is trying to address.

visibility reasoning
perspective reasoning
unanswerable VQA
vision-language models
abstention
Innovation

Methods, ideas, or system contributions that make the work stand out.

visibility reasoning
minimal-edit benchmark
abstention-aware evaluation
perspective reasoning
controlled perturbation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
N
Neil Tripathi
New York University