🤖 AI Summary
This study addresses the limited reasoning capability of large language models (LLMs) under noisy visual evidence caused by the absence of explicit uncertainty signals. To this end, it introduces the VisualNoiseQA benchmark alongside an automated evaluation framework. Methodologically, this work pioneers modeling vision-language models (VLMs) as stochastic sensors and leverages self-consistency techniques to generate calibrated uncertainty measures, which drive LLMs to actively query information for robust visual reasoning. Experiments evaluating multiple LLMs across one thousand complex scenarios reveal significant disparities in their capacity to exploit uncertainty signals for noise-robust inference. Ultimately, this research establishes a scalable evaluation paradigm for multimodal active reasoning.
📝 Abstract
Real-world reasoning rarely reduces to static question answering: agents must actively gather information from tools and sensors that are often noisy and unreliable. Yet most existing active reasoning benchmarks assume that environmental feedback is trustworthy, or introduce noise without exposing an explicit, calibrated uncertainty signal, leaving open how LLMs should reason when the evidence itself is uncertain. We introduce VisualNoiseQA, a novel benchmark for active reasoning under noisy visual feedback. A text-only LLM must solve VQA problems by iteratively querying a fixed, off-the-shelf VLM treated as a stochastic visual sensor. For each query, we draw multiple samples and expose an empirical uncertainty signal via self-consistency, enabling the reasoner to probe from different angles and decide what to ask next and when to stop. Our construction is automatic and scalable: starting from diverse VQA sources and two noisy VLMs, we retain only questions where the sensor is inconsistent yet human-solvable. We evaluate multiple LLM reasoners on 1,000 instances spanning perception, chart understanding, and knowledge-intensive reasoning. VisualNoiseQA thus provides a controlled playground to study how different LLMs exploit uncertainty signals for robust reasoning.