🤖 AI Summary
This study addresses the challenge in zero-shot visual entailment where multi-claim hypotheses complicate reasoning, resulting in performance significantly inferior to fine-tuned models. To this end, we propose a zero-shot inference framework based on atomic fact decomposition and complementary error learning. Built upon frozen vision-language models, the method employs a context-preserving atomic decomposition strategy to break complex hypotheses into atomic propositions, and trains a lightweight classifier to replace traditional majority voting mechanisms for filtering reliable predictions. Experimental results demonstrate that the proposed approach achieves an accuracy of 0.803 on SNLI-VE, accurately localizing visual evidence without fine-tuning and substantially narrowing the performance gap between zero-shot and fine-tuned systems.
📝 Abstract
Visual entailment (VE) asks whether an image supports, contradicts, or leaves undecided a textual hypothesis. Strong results come from fine-tuning large vision-language models on labelled data, while zero-shot and hybrid approaches remain far behind. A VE hypothesis often bundles several visual claims, yet existing zero-shot methods reason over it as a single unit. We propose Atomic Visual Entailment (AVE), which decomposes the hypothesis into atomic facts, produces candidate predictions from both the full hypothesis and its facts using frozen vision-language models, and predicts the final label with a lightweight classifier trained only on how those candidates behave. We find that decomposition helps only when the hypothesis context is preserved: judging facts in isolation is worse than not decomposing at all. Full-hypothesis and atomic prediction make complementary errors, and learning which to trust recovers far more of that complementarity than majority voting, reaching 0.803 test accuracy on SNLI-VE without fine-tuning any vision-language model. AVE also localises the visual evidence behind its prediction without region-level supervision. These results suggest that learning which candidate prediction to trust can close much of the gap to fine-tuned systems, offering a practical alternative where fine-tuning a vision-language model directly would need more labelled data or compute than is available.