🤖 AI Summary
This study addresses the limited comprehension capability of vision-language models in question-answering tasks involving negation clauses. To this end, we propose Skeleton-guided Strategy Planning (SSP), a training-free approach that extracts the structural skeleton of questions and retrieves analogous cases to automatically generate answering strategies via in-context learning, thereby enhancing the model's reasoning over negation logic without parameter updates. By integrating prompt engineering with schema retrieval techniques, SSP effectively improves the parsing of negation semantics. Experimental results demonstrate that SSP achieves state-of-the-art performance across multiple negation VQA benchmarks, exhibiting both computational efficiency and strong generalization capabilities.
📝 Abstract
Despite the strong performance of Vision-Language Models (VLMs) on a wide range of visual question answering (VQA) tasks, these models consistently struggle to understand negation and produce incorrect answers when questions involve negated clauses. To address this limitation, we propose Skeleton-and-Strategy Prompting (\textbf{SSP}), a training-free, in-context learning method that improves VLM negation understanding capabilities without any parameter updates. Given a negation question, our method first abstracts the underlying question structure into a skeleton, retrieves a small set of same-skeleton questions from a lightweight question pool, then prompts the VLM to analyze their shared negation pattern and synthesize a single-sentence answering strategy. The skeleton and strategy are prepended to the test sample to guide the model correctly tackle the negation problems. Experiments on multiple negation VQA benchmarks show that SSP achieves state-of-the-art performance on negation-focused VQA tasks while remaining computationally efficient.