🤖 AI Summary
This study addresses the susceptibility of vision-language models to hallucination and the unreliability of existing single-perspective self-verification methods by proposing MOTIVE, a multi-perspective complementary verification framework. Specifically, MOTIVE evaluates candidate answers from diverse perspectives and learns correctness-aligned reliability scores. Based on these scores, it introduces a dynamic acceptance criterion and a history-informed rethinking strategy, enabling efficient self-correcting reasoning without external judges. Experimental results demonstrate that MOTIVE significantly outperforms strong baselines across multiple multimodal benchmarks. Furthermore, it effectively enhances verification stability and reduces redundant reasoning iterations, achieving more reliable and efficient self-verification for vision-language models.
📝 Abstract
Vision-language models (VLMs) have achieved strong performance in multimodal reasoning, yet they remain prone to generating plausible but incorrect answers. Self-verification offers a practical way to improve answer reliability without relying on external judges, but existing methods typically depend on a single verification criterion or fixed prompt, resulting in incomplete and unstable reliability estimates. We first systematically analyze how verifier capability and prompt design affect verification performance. Our findings show that stronger verifiers provide more reliable judgments, while verification performance is highly sensitive to prompt choice, with no single prompt consistently dominating across tasks. Guided by these findings, we propose \texttt{MOTIVE}, a \textbf{M}ulti-View Self-Verificati\textbf{O}n wi\textbf{T}h Rel\textbf{I}ability-Guided Selecti\textbf{VE} Rethinking framework for reliable multimodal reasoning. \texttt{MOTIVE} evaluates each candidate answer from complementary verification perspectives and learns a correctness-aligned reliability score through correctness-grounded multi-view verification learning. During inference, this score governs an accept-or-rethink decision, allowing reliable answers to be returned directly while uncertain ones trigger history-guided rethinking. Extensive experiments across diverse multimodal benchmarks and VLM backbones demonstrate that \texttt{MOTIVE} consistently outperforms strong self-verification and self-correction baselines. Further results show that reliable verification improves accept-or-rethink decisions and reduces unnecessary reasoning turns, enabling more reliable and efficient self-verification without an external judge.